Memory Sidecar / Fleet audit

Same GPU.
More useful agent work.

2 → 16 is our joint optimization target. If your measured baseline is two long-context sessions, test a path toward sixteen—or the highest ceiling that meets your quality and latency requirements. This is a target, not a measured gain or guarantee.

For GPU operators and inference platforms. Start with a sizing scenario, install the sidecar beside your vLLM, and prove the result on your own tasks before buying. Hosted-API users do not control the underlying GPU bill; this calculator is not a savings estimate for them.

Your evaluation inputs

Count sessions, not people: one developer can run several agents. Type an exact count up to 10,000; the slider covers 1–1,000.

Share of potential sessions active at once. A session can make many model requests.

Use your observed high-water mark, including tool output.

A sizing preset, not a compatibility guarantee. Check actual weights and precision.

Use the setting your model server actually runs.

For long resumed histories, compare adaptive lexical first, then HDC. Adaptive caps only incidental recent history at 4,096 estimated tokens; it is not a 4k prompt. Your choice carries into the download.

Illustrative starting rate. Replace with your contracted rate; zero is allowed. Compute values assume 730 hours/month, not your actual bill. A rate you type stays put when you change GPU.

Exclude other workloads. Owned or prepaid cards do not automatically create cash savings.

Pilot fit · heuristic, not production readiness
Sizing your scenario

Recommended next step: Install in shadow on one host.

Modelled resident sessions per replica · full history to 2× scenario

Baseline cards for this load
Upper-scenario cards
Scenario capacity multiple
Modelled deployment · each block is one GPU
Full-history scenario
Upper improvement scenario
Business-case range · estimates, not bills
Compute capacity value before licence
Your assigned fleet · cards, not dollars
Modelled full-history compute
Upper-scenario compute
Potential freed capacity value · not cash
Potential deferred capacity value
Illustrative deployment licence
Licence · monthly equivalent, paid annually

Do not resize a fleet from this estimate. The range runs from no improvement to an illustrative 2× scenario limited by the selected profile's memory model. It is not a measured result, safety bound or confidence interval. Real performance may be worse. Savings require fewer cancellable GPU-hours or a deferred purchase; fixed commitments may only free capacity. Include licence, CPU/RAM, storage and operating costs in cost per accepted task.

Assumptions and evidence labels

Weights are sized from preset resident GB; the model is sharded across the minimum replica that holds them. VRAM × 0.92 minus weights supplies KV memory; an empirical 0.62 factor adjusts it. Capacity scales with available memory divided by history × KV bytes. This excludes transient prefill memory, prefix-cache effects, scheduler limits and task correctness. Presets are not verified serving configurations. The chosen resident ceiling includes bounded output in live use.

Modelled: every result on this page. Target: the conditional 2 → 16 ambition. Measured: only your completed matched pilot with its pinned hardware, workload and acceptance gates. No measured result for your configuration is attached to this calculator.

Create this evaluation →

From sizing to a buying decision.

1Configure and download

Carry these inputs and your selected profile into Memory. Download a runnable source bundle; evaluation is free.

2Install beside vLLM

Linux + Docker. Provide the local model address and key. The quickstart must pass a real Chat Completions request before it reports success.

3Open the local dashboard

On the GPU host: http://127.0.0.1:8089. For a remote host, use the SSH tunnel shown in the portal. Status and estimated prompt reduction stay local.

4Validate, then license

Calibrate in shadow. Test live at 1, then 2, 4, 8 and 16 only after passing each stage. Compare accepted tasks, p95 latency and total cost before licensing the actual deployed cards.

What to know before your pilot.

Will it fit my stack?

Memory applies to text-only OpenAI /v1/chat/completions on self-hosted vLLM with its generic chat renderer and an equivalent local /tokenize route. The source bundle does not install your model. Other renderers, multimodal inputs and the Responses API are not covered by this live integration.

What if my existing compaction already works?

Keep it as the baseline. Compare native serving, your current compaction, lexical retrieval and HDC under the same conditions. A shorter prompt is not a win if quality drops or cached native turns are faster.

What does 2 → 16 actually mean?

A customer-specific joint optimization target. First measure your baseline; then test progressively. It is not a promise that every GPU starts at two or that this product safely serves sixteen sessions.

What protects the model?

Every live candidate is counted with the local model tokenizer. Prompt plus bounded output for all generation candidates must fit the resident ceiling. Missing preflight or incompatible usage fails closed. Live admission starts at one request; it does not measure free GPU memory or govern clients that bypass the sidecar.

What leaves our network?

The runtime has no telemetry connection to Engrammatic. Prompts and keys remain on your host. Building downloads a base image and packages unless cached. Keep the dashboard on loopback; it shows bounded aggregates, not conversation text.

Can we sell a premium agent tier?

That is the design-partner opportunity: combine your GPU, model, private operations and support with a validated memory policy. An annual software licence does not grant OEM or resale rights; agree commercial terms with us before selling a bundle.

How is it priced?

Per GPU card in each model replica served through the sidecar, charged in advance. Two terms: a 90-day evaluation licence at a quarter of the card's annual rate, which does not auto-renew and is credited in full against an annual licence taken within 30 days of it ending; or the annual licence, billed annually in advance and auto-renewed. You declare the deployment count after testing; changing an active quantity is prorated by Stripe. No usage fee. Cancel through the billing portal before renewal.

What if the pilot fails?

Restore the agent’s original model URL and stop the container. Model weights are unchanged. Keep the native path if it meets your goals better; a failed pilot is not a reason to purchase.