Skip to content

The professional harness for AI coding agents.

Your agents.
Fully equipped.

Run Claude Code, Codex, Kimi, and GLM in one professional harness. Less wasted context. Built-in code intelligence. Smarter task routing.

57.1%
lower AI coding costs25/42 passes in each workflow.
SWE-rebench · 14 tasks · 3 repetitions · interim result

macOS · Windows · Linux · Keep your existing subscriptions.

01 / Bring your favorite agents.02 / Split terminals03 / Token savings04 / Your control
Halv desktop workspace with Claude Code and Codex in three terminal panes, project history, and a savings meter.Halv desktop workspace with Claude Code and Codex in three terminal panes, project history, and a savings meter.
Inside the Halv desktop app

02 / HALV ENGINE

Great agents deserve
a better environment.

The intelligence around your agent matters. Halv brings the tools together, so you can focus on what you’re building.

01 / RTK + Headroom

Token savings, built in

Make every token work harder.

Command output filtering and local context compression keep unnecessary noise out of your agent’s context.

RTK + Headroom
02 / Crux

Your codebase, connected

Give your agent the whole picture.

Callers, references, impact — answered from an index instead of reading the repo a file at a time. Agents ask the index and stop burning context to orient themselves.

Crux
03 / JEV

TASK ROUTING

The right model for the work.

The agent you chat with briefs each task. JEV reads the brief and picks the lane, model and effort: cheaper models for easy work, stronger ones for hard work. Premium models do the work that needs them.

ON BY DEFAULT
03 / THE PROOFSWE-rebench

Same pass count. 57.1% less cost.

Across 42 paired repetitions, both workflows passed 25 runs. Recorded model cost fell from $337.50 to $144.67 with Halv.

Interim result published by Halv: the first 14 of 111 planned tasks. Costs exclude invalid infrastructure attempts and a rolled-back policy experiment.

Inspect cost, tokens, and outcomes.

Inspect cost, tokens, and outcomes.

Halv used 3.49× as many tokens. Model routing reduced recorded cost. Equal aggregate pass counts do not mean identical solved tasks.

METRICCodexCodex + HalvHALV CHANGE
Verifier passes25 / 4225 / 42No change
Total tokens48,755,656170,114,775248.9%
Recorded model cost$337.50$144.67-57.1%
Tokens per correct answer1,950,2266,804,591248.9%
Cost per correct answer$13.50$5.79-57.1%

Same tasks. Different workflows.

Astra medium · Codex 0.155.1 · default service

Halv added JEV-routed Codex workers, Crux, RTK, and Headroom. Worker models and reasoning efforts varied. This comparison measures the combined workflow.

04PRICING

Big savings. Small price.

Halv is built to save you far more than it costs. Get more work from the agent subscriptions you already pay for.

Try it free. Let your savings make the case.

Measure your savings in the app. Keep building with the agent subscriptions you already use.

EVERYTHING INCLUDED

  • Pro mode terminals
  • Four agent CLIs
  • Up to 32 live panes
  • Ghostty on macOS
  • Subagent fan-out
  • Crux code index
  • Command filtering
  • Live savings meter
  • Local OS sandbox
  • Per-project approvals
  • Turn diff & revert
  • macOS · Windows · Linux
ONE PLAN. ALL FEATURES.Unlimited savings for US$10 per month. Your agent subscriptions remain separate.

FREE TO TRY · CANCEL IN TWO CLICKS · WORKS WITH YOUR CLAUDE, CHATGPT, MOONSHOT OR Z.AI PLAN

05FAQ

Questions we get asked.

  • How much does Halv cost, and can I try it free?

    Halv Unlimited costs US$10 per month with no savings cap. It includes all features and a 7-day free trial. A card is required at signup. Cancel during the trial to pay nothing; otherwise, your subscription renews automatically.

  • Does Halv work with Codex and Claude Code?

    Yes. Halv supports Claude Code, Codex, Kimi, and GLM. In Pro mode, each agent runs through its native CLI inside Halv. Sign in with your existing agent account and use the subscriptions you already pay for. API keys are optional for providers that support them.

  • Does Halv run on macOS, Windows, and Linux?

    Yes. Halv supports macOS on Apple silicon and Intel, Windows x64, and Linux x64. Linux downloads include AppImage and deb packages. Halv is a desktop app; download and install it on a supported computer to use the workspace.

  • Does Halv process my code locally?

    Halv runs code indexing and context compression on your computer. Your chosen model provider receives the prompts and context needed to answer your requests. Halv also connects to services for account access, billing, updates, and basic app events. Review agent actions and approvals before commands run or files change.

  • What is Halv?

    Halv is a desktop workspace for AI coding agents. Run Claude Code, Codex, Kimi, and GLM in one app with chat, native terminals, project history, and a live savings meter. Halv combines code navigation, command output filtering, and context compression to reduce wasted tokens while you work.

  • How can I increase my Claude Code usage limit?

    To make your current allowance last longer, reduce unnecessary context and tool output. Halv combines local context compression, code indexing, and command output filtering in one workspace. Anthropic still controls your account's limits.

  • What is the best way to save tokens in Claude Code?

    Start by reducing unnecessary context: keep instructions focused, give Claude relevant files, and use /clear when starting unrelated work. Halv adds code indexing, command output filtering, and local context compression around Claude Code in a desktop workspace. Use its savings meter to measure the difference in your own work. Savings depend on the task and context.

  • How can I increase my Codex usage limit?

    To make your current Codex allowance last longer, keep requests focused and reduce unnecessary context. Halv helps reduce token waste with code indexing, command output filtering, and context compression. OpenAI controls the actual limits.

  • What is the best way to save tokens in Codex?

    Give Codex precise tasks, relevant files, and concise AGENTS.md instructions. Disable unused MCP servers. Measure tokens and cost separately. Halv's latest workflow benchmark used more tokens but recorded 57.1% lower model cost.

  • How does Halv reduce coding agent token usage?

    Halv reduces token usage in three ways. Crux gives agents a code index for navigation and references. RTK filters noisy command output before it enters model context. Halv Engine Headroom compresses model context locally. Together, these tools reduce repeated code exploration and unnecessary context.

  • What is JEV task routing in Halv?

    JEV task routing sends each task to a model that fits it. The agent you chat with acts as coordinator and writes a self-contained brief for each task. JEV, the typesafe/jev routing model from Typesafe, reads the brief and picks a lane (Claude or Codex), a model, and a reasoning effort. Easy work goes to cheaper models, and hard work goes to stronger ones. If a task fails its checks, it escalates to a stronger tier. If your instructions say to implement with Codex or Claude, routing stays in that lane. The task board shows the lane, model, and effort for each routed task. Routing is on by default in Halv 0.6.3. You can turn it off in Settings. Existing explicit preferences are preserved.

  • What evidence supports Halv's cost savings?

    Halv recorded 57.1% lower model cost across 14 SWE-rebench tasks with three repetitions each. Both workflows passed 25 of 42 runs. Halv added routed workers and context tools to Astra medium. The published evidence includes 84 sanitized records and an exclusions ledger.

  • Will every task cost 57.1% less with Halv?

    No. Halv cost more on four of the 14 tasks and used more total tokens. Equal aggregate pass counts hide different solved tasks. This interim Codex result excludes invalid infrastructure attempts and a rolled-back policy experiment. It does not establish Claude Code savings or reduce fixed subscription prices.

  • How do Halv's token savings affect my subscription?

    Halv helps you get more work from your existing agent usage budget. Your provider's fixed subscription price stays the same. Lower token usage can reduce metered API costs, depending on pricing and caching. Compare your measured savings with Halv's price to see whether it pays for your workflow.

  • How is Halv different from a terminal or an MCP server?

    Halv combines coding agents, chat, native terminals, project history, approvals, and savings tracking in one desktop app. Crux is the code-index MCP server within Halv's toolkit. With a standalone CLI, you manage that workspace and those tools yourself. Halv brings them together around the agents you already use.

Put Halv to work on your next task.

Measure your savings in the app. Keep building with the agent subscriptions you already use.

FROM ZERO TO SAVING

01

Download

macOS, Windows or Linux. No API key required.

02

Sign in to your agents

The CLIs you already use, on the plans you already pay for. Logins stay in the CLI.

03

Watch the meter

Savings start on your first prompt, receipt by receipt.

macOS · APPLE SILICON + INTEL

WINDOWS · x64

LINUX · AppImage + deb

7 days free. Then US$10/month. Card required. Cancel during the trial to pay nothing.