More useful agent work from the GPUs you already run. Give long resumed histories a task-relevant working set, then prove the gain on your workload.
Status
- Availability
- Runnable early-access source package. No customer host is changed until its operator installs it.
- Best fit
- Teams operating their own GPUs with large cold or resumed agent histories. Continuously warm native sessions may already be faster.
- Deployment
- One CPU container beside your vLLM. Your client resends the full history; model weights stay unchanged.
- Policy
- Existing 16k dynamic default preserved. Opt-in adaptive profiles use a 32,768-token resident ceiling and a 4,096-estimated-token incidental recent-tail cap. Required records and current work are not tail-capped.
- Joint optimization target
- 2 → 16 on the same GPU—if your measured baseline is two long-context sessions. Test toward sixteen or stop at the highest ceiling that passes your quality and latency gates. This is a conditional target, not a measured gain or guaranteed result.
- What to verify
- Task acceptance, full-task and warm-turn latency, exact prompt counts, cache/recompute work, GPU errors and total cost. Compare native serving, your current compaction, lexical retrieval and HDC.
- Local dashboard
- Installed at http://127.0.0.1:8089 on your GPU host; use an SSH tunnel from your laptop. Profile, live requests, queue, calibration and estimated prompt reduction stay local. Prompt reduction is not cash saved.
- Data boundary
- No outbound telemetry. Adaptive mode retains a bounded in-memory prompt checkpoint per session. Reporting writes only a bounded aggregate snapshot, never prompts or keys. No automatic controller chooses when warm native traffic should be compacted.
Start a controlled pilot
A private workspace configures your next source download. Saving a profile does not change a running host or activate live rewriting.
Evaluation is free; no payment card required. Start in shadow, then run a controlled live comparison. Fleet Audit models a sizing scenario, not a measured saving. License only the GPUs actually deployed after your test passes.
Sign in to configure and download