A coding agent’s answer matters when it passes the checks. The useful cost is everything the workflow spends to get there.
In our latest benchmark checkpoint, Halv recorded less than half the model cost while matching vanilla Codex’s total number of passes.
We gave both workflows the same 14 repository tasks and ran each task three times. That produced 42 paired comparisons, or 84 runs in total. Repository verifiers checked the resulting work.
The result in three numbers
| Vanilla Codex | Codex + Halv | |
|---|---|---|
| Runs that passed | 25/42 | 25/42 |
| Recorded model cost | $337.50 | $144.67 |
| Cost per passing run | $13.50 | $5.79 |
Halv recorded 57.1% lower model cost. The difference was $192.83 across these selected runs. Both cost totals include valid runs that failed their checks.
JEV, which selects the agents, is included in the Halv subscription price. You do not pay a separate fee for that selection. The benchmark totals measure model usage, with the Halv subscription priced separately.
The pass totals match, but the solved tasks differ. Halv gained passes on two tasks and lost passes on another. The result shows an aggregate cost improvement for this sample.
The best example: the same three passes for much less
On ArcadeDB-4455, both workflows passed all three repetitions. Vanilla recorded $81.85 in model cost. Halv recorded $14.97.
That is 81.7% lower cost with 3/3 passes on both sides.
Two other strong results help explain the overall number:
- HugeGraph-3037: 67.1% lower cost; Halv passed 3/3 runs, while vanilla passed 2/3.
- ArcadeDB-4411: 46.1% lower cost; both passed 3/3 runs.
These were the highlights, not the whole story. Halv cost more on four of the 14 tasks. On Perry-3982, it cost less but passed only 1/3 runs, compared with vanilla’s 3/3.
What Halv did differently
Both workflows started with Astra at medium reasoning. Vanilla used that agent to do the work alone.
Halv added a team structure. A cheaper agent investigated the task. JEV selected a worker for the task’s difficulty. That worker could delegate smaller work when useful. Astra reviewed the result.
Halv also supplied code navigation, filtered command output, and context processing through Crux, RTK, and Headroom. This benchmark measured the complete workflow. It did not measure each feature separately.
More tokens can still cost less
Halv used 3.49 times as many total tokens across its coordinator and workers. Its recorded model cost was still lower. Different models and cached requests have different costs.
This result is about the cost of the complete workflow. It does not show that adding agents automatically saves tokens. It also does not reduce the fixed price of a Codex or Claude subscription.
What to keep in mind
This is an interim result published by Halv, covering the first 14 tasks in a 111-task plan. We paused after a complete task at the user’s request. Three repetitions help show variation, but this remains a small sample.
The costs are recorded model estimates, not invoices or the entire benchmark bill. They include all recorded coordinator and worker sessions in the selected runs. They exclude invalid infrastructure attempts, a rolled-back policy experiment, and operational costs. Some excluded attempts have unknown costs.
The result does not guarantee the same savings on your tasks. It also does not establish equal answer quality or Claude Code performance.
Read the technical report
The technical report includes every task, the model configuration, role costs, selection rules, and exclusions. The public evidence includes all 42 pairs and a script that checks the published arithmetic.
For your own workflow, compare both the complete cost and the work that passes your checks. That is the result this benchmark makes visible.
Frequently asked questions
What is the main result?
Halv recorded $144.67 in model cost, compared with $337.50 for vanilla Codex. Both passed 25 of 42 runs across 14 tasks. That is 57.1% lower recorded cost with the same aggregate pass count.
Did every task get cheaper?
No. Halv cost less on ten tasks and more on four. It also passed fewer repetitions on one task. The complete technical report includes every task.
Does lower cost mean fewer tokens?
Not in this benchmark. Halv used 3.49 times as many tokens across its coordinator and workers. The workflow used different models and recorded lower total model cost.
Will my subscription become cheaper?
No. The result measures recorded model usage cost. Fixed subscription prices stay the same, and this checkpoint does not predict every user's savings.