Skip to content

FOR ANYONE WHO HITS LIMITS

Make your plan go
twice as far.

Halv routes every chat through a local engine that strips wasted tokens before they reach the model. Same model, same answers — you just hit your limits later.

7-day free trial

THE APP

The Halv desktop app: a chat answer with its compression receipt, next to a live savings panel that shows tokens saved today.
The Halv desktop app: a chat answer with its compression receipt, next to a live savings panel that shows tokens saved today.

Every answer shows its receipt. The sidebar keeps the running total.

PRO MODE

Real terminals. Built for power users.

Run each agent in its native CLI — inside Halv’s embedded terminal on every platform, or in Ghostty windows on macOS with Ghostty 1.3+. Halv still measures the savings.

Halv Pro mode with multiple agents running in native CLI terminals.
Halv Pro mode with multiple agents running in native CLI terminals.

Native CLIs · embedded everywhere · Ghostty on macOS

HOW IT WORKS

Nothing new to learn.
Half the bill.

Halv sits quietly between you and the model. You work exactly as you do today — it just makes every request cheaper on the way through.

  1. 01

    Sign in with ChatGPT

    One click, the official sign-in flow. No API keys to paste, nothing to configure.

  2. 02

    Work like you already do

    Ask questions, write, or hand Halv a coding task in your repo. Same models, same quality you expect.

  3. 03

    Spend half the tokens

    Halv compresses every exchange in the background. Same answers, and the savings add up.

THE HALV ENGINE

Not just cheaper.
More correct.

Halv runs every request through its own engine: three systems, each measured.

Compression

Rewrites bloated context before it reaches the model.

~64%context reduction

the live demo in the hero is real

Crux code index

A map of your codebase: callers, references, impact — answered from an index instead of reading files.

96% vs 66%correct

−24%cost per answer

Read the benchmark →

Open source, MIT.

Command filtering

Trims noisy command output — test logs, diffs, dependency trees — before it costs you tokens.

test log · 3,104 lines → 41 lines

illustrative example

counted in the live meter

PROOF

Numbers we can defend.

Claim
Correct answers on large codebases
Number
96% vs 66%
How it was measured
50 verifiable questions over Django and SymPy; the same agent with and without the Crux code index
Claim
Tokens per correct answer
Number
−24%
How it was measured
Same benchmark, token-metered through a local proxy
Claim
Budget burned on wrong answers
Number
5% vs 35%
How it was measured
Same benchmark: share of total tokens spent on answers that turned out wrong
Claim
Full-session savings
Number
In progress
How it was measured
halv-bench, a session-shaped benchmark of the complete engine. Results will appear here.
Source
IN PROGRESS

PRICING

7-day free trial. Cancel in two clicks.

Basic is $2/month for ChatGPT Plus; Unlimited is $10/month for $100+ plans, and both add Halv to the plan you already pay for.

FAQ

Questions,
answered.

Still curious? Write to [email protected].

Do I need an OpenAI API key?

No. Sign in with your existing ChatGPT account and Halv uses the plan you already pay for. If you'd rather use an API key, that works too.

How can it use fewer tokens and still give the same answer?

Halv strips what the model was going to ignore — duplicated file context, stale history, boilerplate instructions — and keeps everything that changes the answer. When a request can't be compressed safely, it passes through untouched.

Is my code safe?

Everything runs locally. Halv reads your repo on your machine, and nothing leaves except the model call itself — the same call your client already makes. Shell commands run in an OS sandbox behind approvals you control.

Which platforms are supported?

Halv is available for macOS on Apple silicon and Intel, Windows x64, and Linux x64 as AppImage or deb packages. Both plans include a 7-day free trial.

Is Halv only for programmers?

No. Halv works with any ChatGPT conversation, including questions, writing, and research. The coding tools are there when you want them, and compression saves on everything.

How is this different from the ChatGPT app or Codex?

Same models underneath. Halv adds the compression layer, the project sandbox, and a live savings meter — you keep your workflow and your plan, and spend about half as much of it.

Start saving tokens
for free.

Download Halv for macOS, Windows, or Linux. Sign in, choose your plan, and start your 7-day free trial inside the app.

7-day free trial · cancel anytime during the trial and pay nothing.