LLM Memory Sidecar

Prove the fit on your own workload

controlled-pilot early access

More useful agent work from the GPUs you already run. Give long resumed histories a task-relevant working set, then prove the gain on your workload.

Status

Availability
Runnable early-access source package. No customer host is changed until its operator installs it.
Best fit
Teams operating their own GPUs with large cold or resumed agent histories. Continuously warm native sessions may already be faster.
Deployment
One CPU container beside your vLLM. Your client resends the full history; model weights stay unchanged.
Policy
Existing 16k dynamic default preserved. Opt-in adaptive profiles use a 32,768-token resident ceiling and a 4,096-estimated-token incidental recent-tail cap. Required records and current work are not tail-capped.
Joint optimization target
2 → 16 on the same GPU—if your measured baseline is two long-context sessions. Test toward sixteen or stop at the highest ceiling that passes your quality and latency gates. This is a conditional target, not a measured gain or guaranteed result.
What to verify
Task acceptance, full-task and warm-turn latency, exact prompt counts, cache/recompute work, GPU errors and total cost. Compare native serving, your current compaction, lexical retrieval and HDC.
Local dashboard
Installed at http://127.0.0.1:8089 on your GPU host; use an SSH tunnel from your laptop. Profile, live requests, queue, calibration and estimated prompt reduction stay local. Prompt reduction is not cash saved.
Data boundary
No outbound telemetry. Adaptive mode retains a bounded in-memory prompt checkpoint per session. Reporting writes only a bounded aggregate snapshot, never prompts or keys. No automatic controller chooses when warm native traffic should be compacted.

Start a controlled pilot

A private workspace configures your next source download. Saving a profile does not change a running host or activate live rewriting.

Evaluation is free; no payment card required. Start in shadow, then run a controlled live comparison. Fleet Audit models a sizing scenario, not a measured saving. License only the GPUs actually deployed after your test passes.

Sign in to configure and download