2 → 16 is our joint optimization target. If your measured baseline is two long-context sessions, test a path toward sixteen—or the highest ceiling that meets your quality and latency requirements. This is a target, not a measured gain or guarantee.
For GPU operators and inference platforms. Start with a sizing scenario, install the sidecar beside your vLLM, and prove the result on your own tasks before buying. Hosted-API users do not control the underlying GPU bill; this calculator is not a savings estimate for them.
Count sessions, not people: one developer can run several agents. Type an exact count up to 10,000; the slider covers 1–1,000.
Share of potential sessions active at once. A session can make many model requests.
Use your observed high-water mark, including tool output.
A sizing preset, not a compatibility guarantee. Check actual weights and precision.
Use the setting your model server actually runs.
For long resumed histories, compare adaptive lexical first, then HDC. Adaptive caps only incidental recent history at 4,096 estimated tokens; it is not a 4k prompt. Your choice carries into the download.
Illustrative starting rate. Replace with your contracted rate; zero is allowed. Compute values assume 730 hours/month, not your actual bill. A rate you type stays put when you change GPU.
preset default /h —
Exclude other workloads. Owned or prepaid cards do not automatically create cash savings.
Recommended next step: Install in shadow on one host.
Do not resize a fleet from this estimate. The range runs from no improvement to an illustrative 2× scenario limited by the selected profile's memory model. It is not a measured result, safety bound or confidence interval. Real performance may be worse. Savings require fewer cancellable GPU-hours or a deferred purchase; fixed commitments may only free capacity. Include licence, CPU/RAM, storage and operating costs in cost per accepted task.
Weights are sized from preset resident GB; the model is sharded across the minimum replica that holds them. VRAM × 0.92 minus weights supplies KV memory; an empirical 0.62 factor adjusts it. Capacity scales with available memory divided by history × KV bytes. This excludes transient prefill memory, prefix-cache effects, scheduler limits and task correctness. Presets are not verified serving configurations. The chosen resident ceiling includes bounded output in live use.
Modelled: every result on this page. Target: the conditional 2 → 16 ambition. Measured: only your completed matched pilot with its pinned hardware, workload and acceptance gates. No measured result for your configuration is attached to this calculator.
We will size your actual fleet and tell you which parts we have measured and which we are modelling.
Carry these inputs and your selected profile into Memory. Download a runnable source bundle; evaluation is free.
Linux + Docker. Provide the local model address and key. The quickstart must pass a real Chat Completions request before it reports success.
On the GPU host: http://127.0.0.1:8089. For a remote host, use the SSH tunnel shown in the portal. Status and estimated prompt reduction stay local.
Calibrate in shadow. Test live at 1, then 2, 4, 8 and 16 only after passing each stage. Compare accepted tasks, p95 latency and total cost before licensing the actual deployed cards.
Memory applies to text-only OpenAI /v1/chat/completions on self-hosted vLLM with its generic chat renderer and an equivalent local /tokenize route. The source bundle does not install your model. Other renderers, multimodal inputs and the Responses API are not covered by this live integration.
Keep it as the baseline. Compare native serving, your current compaction, lexical retrieval and HDC under the same conditions. A shorter prompt is not a win if quality drops or cached native turns are faster.
A customer-specific joint optimization target. First measure your baseline; then test progressively. It is not a promise that every GPU starts at two or that this product safely serves sixteen sessions.
Every live candidate is counted with the local model tokenizer. Prompt plus bounded output for all generation candidates must fit the resident ceiling. Missing preflight or incompatible usage fails closed. Live admission starts at one request; it does not measure free GPU memory or govern clients that bypass the sidecar.
The runtime has no telemetry connection to Engrammatic. Prompts and keys remain on your host. Building downloads a base image and packages unless cached. Keep the dashboard on loopback; it shows bounded aggregates, not conversation text.
That is the design-partner opportunity: combine your GPU, model, private operations and support with a validated memory policy. An annual software licence does not grant OEM or resale rights; agree commercial terms with us before selling a bundle.
Per GPU card in each model replica served through the sidecar, charged in advance. Two terms: a 90-day evaluation licence at a quarter of the card's annual rate, which does not auto-renew and is credited in full against an annual licence taken within 30 days of it ending; or the annual licence, billed annually in advance and auto-renewed. You declare the deployment count after testing; changing an active quantity is prorated by Stripe. No usage fee. Cancel through the billing portal before renewal.
Restore the agent’s original model URL and stop the container. Model weights are unchanged. Keep the native path if it meets your goals better; a failed pilot is not a reason to purchase.