# Halv: product facts and FAQ > The desktop workspace for AI coding agents. Website: https://halv.ai/ ## Plans - Unlimited: US$10 per month. Every feature, with no cap on savings. Give your agents more room to work, across projects and sessions, all month long. Halv Unlimited costs US$10 per month with no savings cap. It includes all features and a 7-day free trial. A card is required at signup. Cancel during the trial to pay nothing; otherwise, your subscription renews automatically. ## Task routing JEV agent selection is included in your Halv subscription, with no separate routing fee. Halv 0.6.0 adds built-in task routing powered by JEV, the typesafe/jev routing model from Typesafe. The agent you chat with acts as coordinator and writes a self-contained brief for each task. JEV reads the brief and picks a lane (Claude or Codex), a model, and a reasoning effort. - Easy work goes to cheaper models, for example Claude Haiku 4.5, Claude Sonnet 5, or GPT-6 Luna. - Hard work goes to stronger models, for example Claude Opus 5.5 or GPT-6 Sol. The hardest work stays with the coordinator. - If a worker fails its checks, the task escalates to a stronger tier. - Standing instructions apply. If your setup says to implement with Codex or Claude, routing stays in that lane. - The task board shows the lane, model, and effort for each routed task. - Routing is on by default in Halv 0.6.3. You can turn it off in Settings. Existing explicit preferences are preserved. Task routing is separate from the request path. Compression, the Crux code index, and RTK command filtering still run in front of every request. The latest benchmark evaluates an experimental JEV hierarchy with the complete Halv toolkit. It does not isolate routing from other components. ## Published benchmark Across 14 SWE-rebench tasks with 3 repetitions each, Halv recorded 57.1% lower model cost. Both workflows passed 25/42 runs. Halv cost $144.67; vanilla cost $337.50. Halv used 3.49 times as many total tokens. Recorded cost per correct answer was 57.1% lower. Recorded benchmark cost is not a reduction in a fixed subscription bill. Both coordinators used openai/gpt-6-astra, medium reasoning, default service, and Codex 0.155.1. Halv added Luna and Sol workers with different reasoning efforts. Each pair used the same task and repository verifier. The Halv arm used JEV hierarchy routing, Crux, RTK, and Halv Engine Headroom together. The result measures the combined workflow. The first 14 tasks in the fixed 111-task plan, with three repetitions per task. Select the first valid run for each task, arm, and repetition. Keep valid verifier failures. Pause after Pulsar-25793 at the user's request. Halv published this interim workflow comparison. It covers the first 14 of 111 planned tasks. Equal aggregate pass counts do not imply identical solved tasks. The result does not establish Claude Code savings or independent endorsement by SWE-rebench. Infrastructure-invalid attempts and a rolled-back policy experiment are excluded. Some excluded costs are unknown. Costs and tokens include all coordinator and worker sessions in the selected runs, including valid verifier failures. A correct answer means the repository verifier accepted the result. Recorded cost is a token-based estimate, not an invoice. The public script checks arithmetic and file integrity; it does not rerun tasks or authenticate provider billing. - [Overview](https://halv.ai/blog/halv-57-percent-lower-model-cost) - [Technical report and methodology](https://halv.ai/blog/halv-swe-rebench-astra-42-pairs) - [All 84 selected run records](https://halv.ai/evidence/swe-rebench-astra-42-pairs/) - [Machine-readable results](https://halv.ai/evidence/swe-rebench-astra-42-pairs/selected-pairs.json) - [Evidence archive](https://halv.ai/evidence/swe-rebench-astra-42-pairs/halv-swe-rebench-astra-42-pairs-evidence.tar.gz) ## Frequently asked questions ### What is Halv? Halv is a desktop workspace for AI coding agents. Run Claude Code, Codex, Kimi, and GLM in one app with chat, native terminals, project history, and a live savings meter. Halv combines code navigation, command output filtering, and context compression to reduce wasted tokens while you work. [This answer on the website](https://halv.ai/#faq-what-is-halv) - [See the workspace](https://halv.ai/#workspace) ### How can I increase my Claude Code usage limit? To make your current allowance last longer, reduce unnecessary context and tool output. Halv combines local context compression, code indexing, and command output filtering in one workspace. Anthropic still controls your account's limits. [This answer on the website](https://halv.ai/#faq-increase-claude-code-usage-limit) - [Read the guide](https://halv.ai/blog/claude-usage-limits-explained) - [Official Claude Code guidance](https://code.claude.com/docs/en/costs#reduce-token-usage) ### What is the best way to save tokens in Claude Code? Start by reducing unnecessary context: keep instructions focused, give Claude relevant files, and use /clear when starting unrelated work. Halv adds code indexing, command output filtering, and local context compression around Claude Code in a desktop workspace. Use its savings meter to measure the difference in your own work. Savings depend on the task and context. [This answer on the website](https://halv.ai/#faq-save-tokens-claude-code) - [Read the guide](https://halv.ai/blog/how-to-save-tokens-in-claude) - [Official Claude Code guidance](https://code.claude.com/docs/en/costs#reduce-token-usage) ### How can I increase my Codex usage limit? To make your current Codex allowance last longer, keep requests focused and reduce unnecessary context. Halv helps reduce token waste with code indexing, command output filtering, and context compression. OpenAI controls the actual limits. [This answer on the website](https://halv.ai/#faq-increase-codex-usage-limit) - [Read the guide](https://halv.ai/blog/codex-usage-limits-explained) - [Official Codex guidance](https://developers.openai.com/codex/pricing/#what-can-i-do-to-make-my-usage-limits-last-longer) ### What is the best way to save tokens in Codex? Give Codex precise tasks, relevant files, and concise AGENTS.md instructions. Disable unused MCP servers. Measure tokens and cost separately. Halv's latest workflow benchmark used more tokens but recorded 57.1% lower model cost. [This answer on the website](https://halv.ai/#faq-save-tokens-codex) - [Read the guide](https://halv.ai/blog/how-to-save-tokens-in-chatgpt) - [Read the benchmark](https://halv.ai/blog/halv-57-percent-lower-model-cost) - [Official Codex guidance](https://developers.openai.com/codex/pricing/#what-can-i-do-to-make-my-usage-limits-last-longer) ### Does Halv work with Codex and Claude Code? Yes. Halv supports Claude Code, Codex, Kimi, and GLM. In Pro mode, each agent runs through its native CLI inside Halv. Sign in with your existing agent account and use the subscriptions you already pay for. API keys are optional for providers that support them. [This answer on the website](https://halv.ai/#faq-supported-agents) ### How does Halv reduce coding agent token usage? Halv reduces token usage in three ways. Crux gives agents a code index for navigation and references. RTK filters noisy command output before it enters model context. Halv Engine Headroom compresses model context locally. Together, these tools reduce repeated code exploration and unnecessary context. [This answer on the website](https://halv.ai/#faq-token-savings) - [HOW IT WORKS](https://halv.ai/#engine) ### What is JEV task routing in Halv? JEV task routing sends each task to a model that fits it. The agent you chat with acts as coordinator and writes a self-contained brief for each task. JEV, the typesafe/jev routing model from Typesafe, reads the brief and picks a lane (Claude or Codex), a model, and a reasoning effort. Easy work goes to cheaper models, and hard work goes to stronger ones. If a task fails its checks, it escalates to a stronger tier. If your instructions say to implement with Codex or Claude, routing stays in that lane. The task board shows the lane, model, and effort for each routed task. Routing is on by default in Halv 0.6.3. You can turn it off in Settings. Existing explicit preferences are preserved. [This answer on the website](https://halv.ai/#faq-task-routing) - [HOW IT WORKS](https://halv.ai/#engine) ### What evidence supports Halv's cost savings? Halv recorded 57.1% lower model cost across 14 SWE-rebench tasks with three repetitions each. Both workflows passed 25 of 42 runs. Halv added routed workers and context tools to Astra medium. The published evidence includes 84 sanitized records and an exclusions ledger. [This answer on the website](https://halv.ai/#faq-benchmark-evidence) - [Read the benchmark](https://halv.ai/blog/halv-57-percent-lower-model-cost) - [Explore all 42 pairs](https://halv.ai/evidence/swe-rebench-astra-42-pairs/) ### Will every task cost 57.1% less with Halv? No. Halv cost more on four of the 14 tasks and used more total tokens. Equal aggregate pass counts hide different solved tasks. This interim Codex result excludes invalid infrastructure attempts and a rolled-back policy experiment. It does not establish Claude Code savings or reduce fixed subscription prices. [This answer on the website](https://halv.ai/#faq-savings-and-quality) ### How do Halv's token savings affect my subscription? Halv helps you get more work from your existing agent usage budget. Your provider's fixed subscription price stays the same. Lower token usage can reduce metered API costs, depending on pricing and caching. Compare your measured savings with Halv's price to see whether it pays for your workflow. [This answer on the website](https://halv.ai/#faq-subscription-value) - [PRICING](https://halv.ai/#pricing) ### How is Halv different from a terminal or an MCP server? Halv combines coding agents, chat, native terminals, project history, approvals, and savings tracking in one desktop app. Crux is the code-index MCP server within Halv's toolkit. With a standalone CLI, you manage that workspace and those tools yourself. Halv brings them together around the agents you already use. [This answer on the website](https://halv.ai/#faq-cli-and-mcp) ### How much does Halv cost, and can I try it free? Halv Unlimited costs US$10 per month with no savings cap. It includes all features and a 7-day free trial. A card is required at signup. Cancel during the trial to pay nothing; otherwise, your subscription renews automatically. [This answer on the website](https://halv.ai/#faq-pricing-and-trial) - [Download Halv](https://halv.ai/download) ### Does Halv process my code locally? Halv runs code indexing and context compression on your computer. Your chosen model provider receives the prompts and context needed to answer your requests. Halv also connects to services for account access, billing, updates, and basic app events. Review agent actions and approvals before commands run or files change. [This answer on the website](https://halv.ai/#faq-privacy) - [Privacy](https://halv.ai/privacy) ### Does Halv run on macOS, Windows, and Linux? Yes. Halv supports macOS on Apple silicon and Intel, Windows x64, and Linux x64. Linux downloads include AppImage and deb packages. Halv is a desktop app; download and install it on a supported computer to use the workspace. [This answer on the website](https://halv.ai/#faq-platforms) - [All platforms](https://halv.ai/download) ## Official links - [Download Halv](https://halv.ai/download) - [Privacy](https://halv.ai/privacy) - [Terms](https://halv.ai/terms) - [Contact](mailto:hello@halv.ai)