Cut your AI token bill in half.
Agents re-send everything they've seen on every turn — and you pay for all of it. Stark compiles what your tools produce into what your model actually needs: 44–78% less provider input, zero quality loss, and every original byte recoverable on demand.
Proof on your own machine
The meter is built in. Verify us in one session.
Every wrapped command records what came in raw and what was sent. No dashboards to trust — one command, your workload, your numbers.
$ stark savings wrapped-commands: 42 (compiled: 17, passed-through: 25) tool-output tokens est: 184,220 raw -> 21,904 sent reduction: 88.1% $ stark expand 77c1b362… exact original restored, byte for byte
How it works
Wrap. Compile. Expand.
You already do context engineering by hand — trimming logs, re-pasting the one error that matters. Stark does it automatically, deterministically, and keeps receipts.
1Evidence stays local
Stark intercepts noisy read-only tool output — tests, searches, logs, large JSON — and archives the exact bytes in a content-addressed ledger on your disk. No network. No model calls. No telemetry.
2The model sees what matters
A deterministic compiler ranks errors, identifiers, counts, and task-relevant records into a hard token budget — even when the answer sits on page forty. Dense content passes through untouched.
3Nothing is ever lost
Every compiled view carries an object ID. One command restores the original byte-for-byte — full detail exactly when it matters.
The safety contract: diffs, short outputs, and dense content pass through unchanged; errors and exact locations are always protected; a quality regression fails our evals even when tokens improve.
Where it runs
Every serious coding agent.
| Host | Status | Integration |
|---|---|---|
| Claude Code | Available | Native plugin — automatic hooks, zero workflow change |
| Codex | Available | Native plugin — automatic hooks, zero workflow change |
| Cursor | Available | Universal shell adapter (stark shim) — quality proven in cross-host evals |
| Grok CLI | Available | Universal shell adapter (stark shim) — quality proven in cross-host evals |
| Gemini CLI | Available | Universal shell adapter (stark shim) |
| Production AI apps | Reference live | Library at your model-facing seam — staged rollout, per-turn counterfactuals |
One compiler, two doors: hosts with hook systems get native plugins; every other agent that runs shell commands gets the shim — activated only inside the agent's environment, never your own terminal.
Two bills, one fix
Your agents burn tokens. So does your product.
The same waste shows up twice: in the coding agents your team runs all day, and in every API call your application makes. Most teams cope by downgrading to weaker models. Stark cuts the context instead — you keep the max models and lose the bill.
While you build
Agents re-read their whole history every turn. Stark stretches every plan and allowance 1.5–2× on context-heavy work — the meter proves it in your first session. Start free →
While you ship
Your product pays the same tax on every large tool result it sends a model. The same compiler wires into your model-facing seam with staged rollout and replayable savings. Enterprise →
Who feels it
Engineers stop hitting limits mid-task. Platform leads get fleet-level policy and savings rollups. CTOs and CFOs get an AI line item that shrinks with an audit trail behind every dollar.
Why now
Token bills are the new cloud bills.
Prices per token keep falling. Tokens per task rise faster — because agent workflows re-read their whole history on every step.
Measured, not promised
Graded on required facts. Quoted from the invoice.
| Host | Provider input | Reduction | Answer quality |
|---|---|---|---|
| Claude Code | 365,350 → 121,176 | −66.8% | 4/4 non-inferior · one case improved |
| Codex | 280,365 → 80,778 | −71.2% | 4/4 non-inferior |
| Codex · CRM JSON | 101,423 → 44,734 | −55.8% | 100% facts · 2.27× work per token |
| Cursor · Grok | cross-host battery | — | no semantic quality loss in any cell |
| vs. first-page truncation | same token budget | 13/13 facts | truncation kept 5/13 |
The honesty note: a 99% payload cut lands as 44–78% on real invoices because providers add fixed prompt overhead. We quote the invoice number. Deterministic suites hold 90/90 required facts across 3.9M raw tokens, and dense content passes through untouched.
Everything it does
Everything you need to save.
One runtime for the whole token problem — measure, compile, recover, prove.
Automatic wrapping
Audited read-only commands compile with zero workflow change
Deterministic compiler
No model calls — same input, same output, forever testable
Evidence ledger
Exact bytes archived locally, content-addressed by SHA-256
Byte-exact expand
Restore any original with one command when detail matters
Savings meter
Raw vs sent tokens per command, month recaps, milestones
Fair by default
Unused months refund themselves; cancel falls back to pass-through
Native plugins
Claude Code and Codex hooks — install once, forget it
Universal shim
Cursor, Grok, Gemini, and any agent that runs commands
Production library
The same compiler at your product's model-facing seam
Proof-of-value mode
Measure on live traffic without changing what users see
Safety contract
Errors, diffs, and dense content always protected or passed through
Sovereign ready
No network path, no telemetry — air-gapped is native