Crux in real sessions: 90% correct, 47% cheaper per answer
The follow-up benchmark: Crux in real multi-question coding sessions on django and sympy. 90% correct answers versus grep's 58% — at 47% less cost per correct answer.
Read postHALV BLOG
News about Halv and notes from the team building it.
The follow-up benchmark: Crux in real multi-question coding sessions on django and sympy. 90% correct answers versus grep's 58% — at 47% less cost per correct answer.
Read postMethodology and full data for the session-shaped Crux benchmark, the 2.3x token blowup it exposed on sympy, and the Crux 0.6.2 output redesign that turned it into a 59% saving.
Read postWe benchmarked an AI coding agent with and without Crux, our code-index MCP server, in Codex on django and sympy: 45% more correct answers, 24% less cost per answer.
Read postMethodology and full data for the Crux benchmark: an SCIP code-index MCP server vs grep in Codex CLI, with token-metered paired runs, ablations, and the tuning trail.
Read post