Compression
Rewrites bloated context before it reaches the model.
~64%context reduction
the live demo in the hero is real
FOR ANYONE WHO HITS LIMITS
Halv routes every chat through a local engine that strips wasted tokens before they reach the model. Same model, same answers — you just hit your limits later.
7-day free trial
LIVE SAVINGS
today
612.4K
TOKENS SAVED
PER REPLY · TODAY
PLAN WINDOW
Illustrative numbers. Your meter starts at zero.
THE APP


Every answer shows its receipt. The sidebar keeps the running total.
PRO MODE
Run each agent in its native CLI — inside Halv’s embedded terminal on every platform, or in Ghostty windows on macOS with Ghostty 1.3+. Halv still measures the savings.


Native CLIs · embedded everywhere · Ghostty on macOS
HOW IT WORKS
Halv sits quietly between you and the model. You work exactly as you do today — it just makes every request cheaper on the way through.
One click, the official sign-in flow. No API keys to paste, nothing to configure.
Ask questions, write, or hand Halv a coding task in your repo. Same models, same quality you expect.
Halv compresses every exchange in the background. Same answers, and the savings add up.
THE HALV ENGINE
Halv runs every request through its own engine: three systems, each measured.
Rewrites bloated context before it reaches the model.
~64%context reduction
the live demo in the hero is real
A map of your codebase: callers, references, impact — answered from an index instead of reading files.
96% vs 66%correct
−24%cost per answer
Open source, MIT.
Trims noisy command output — test logs, diffs, dependency trees — before it costs you tokens.
test log · 3,104 lines → 41 lines
illustrative example
counted in the live meter
Halve your token spend from the next prompt.
PROOF
PRICING
Basic is $2/month for ChatGPT Plus; Unlimited is $10/month for $100+ plans, and both add Halv to the plan you already pay for.
No. Sign in with your existing ChatGPT account and Halv uses the plan you already pay for. If you'd rather use an API key, that works too.
Halv strips what the model was going to ignore — duplicated file context, stale history, boilerplate instructions — and keeps everything that changes the answer. When a request can't be compressed safely, it passes through untouched.
Everything runs locally. Halv reads your repo on your machine, and nothing leaves except the model call itself — the same call your client already makes. Shell commands run in an OS sandbox behind approvals you control.
Halv is available for macOS on Apple silicon and Intel, Windows x64, and Linux x64 as AppImage or deb packages. Both plans include a 7-day free trial.
No. Halv works with any ChatGPT conversation, including questions, writing, and research. The coding tools are there when you want them, and compression saves on everything.
Same models underneath. Halv adds the compression layer, the project sandbox, and a live savings meter — you keep your workflow and your plan, and spend about half as much of it.
Download Halv for macOS, Windows, or Linux. Sign in, choose your plan, and start your 7-day free trial inside the app.
7-day free trial · cancel anytime during the trial and pay nothing.