Benzi Product comes out on top (9.6/10), followed by Bare CC (8.8), Benzi MCP (8.4), and CodeGraph (8.1).
Overview & comparative ranking
| Rank | Arm | Score | Core answer quality profile |
|---|---|---|---|
| 1 | Benzi Product | 9.6 | Unmatched technical depth, architectural insights, edge-case coverage, and structural visualizations (flowcharts/tables). |
| 2 | Bare CC | 8.8 | Exceptionally precise, concise line-by-line tracing, high citation accuracy, and immediate zero-fluff code mechanics. |
| 3 | Benzi MCP | 8.4 | Strong overall structure, clear stage-by-stage breakdowns, and reliable coverage of core mechanisms. |
| 4 | CodeGraph | 8.1 | Highly readable, great code snippets and summaries, but slightly less granular on deep runtime invariants. |
Benzi Product comes out on top (9.1/10), followed by Bare CC (8.9), Benzi MCP (8.5), and CodeGraph (8.2).
Overall average and ranking
| Rank | Arm | Rating | Why |
|---|---|---|---|
| 1 | Benzi Product | 9.1 / 10 | Most complete and precise overall. Usually gives the deepest mechanism-level explanation, strong line refs, and useful architectural caveats. If you ignore its meta preambles, the actual answers are the best. |
| 2 | Bare CC | 8.9 / 10 | Very clean, direct, accurate, and well-structured. Often nearly as good as Benzi Product with less verbosity. Minor factual slip: in Gin it says abortIndex is 127, but it's math.MaxInt8 >> 1 = 63. |
| 3 | Benzi MCP | 8.5 / 10 | Accurate, well-organized, good line refs and flow descriptions. Usually slightly less exhaustive than the top two, but rarely wrong. |
| 4 | CodeGraph | 8.2 / 10 | Concise and accurate. Good for quick, readable explanations, but often misses some depth/nuance that the top arms include. |
Benzi Product comes out on top (avg 8.7). CodeGraph and Benzi MCP tie at 7.8, and Bare CC is last at 7.6.
Score by repo
| Repo | Bare CC | CodeGraph | Benzi MCP | Benzi Product |
|---|---|---|---|---|
| gin | 6.5 | 7.5 | 7.5 | 9 |
| alamofire | 9 | 8.5 | 8 | 7 |
| okhttp | 7 | 8 | 7.5 | 9.5 |
| excalidraw | 7.5 | 7.5 | 8 | 9 |
| tokio | 8.5 | 7.5 | 8 | 9 |
| django | 7 | 8 | 7.5 | 8.5 |
| Average | 7.6 | 7.8 | 7.8 | 8.7 |
Benzi Product comes out on top (9.73/10 avg), followed by Benzi MCP (9.15), CodeGraph (9.10), and Bare CC (8.97).
Per-repo ranking
| Repo | 🥇 | 🥈 | 🥉 | 4th |
|---|---|---|---|---|
| Gin | Benzi Product — 9.8 | Benzi MCP — 7.7 | CodeGraph — 7.6 | Bare CC — 7.4 |
| Alamofire | Bare CC — 9.6 | Benzi Product — 9.5 | Benzi MCP — 9.4 | CodeGraph — 9.0 |
| OkHttp | Benzi Product — 9.8 | CodeGraph — 9.7 | Bare CC — 9.3 | Benzi MCP — 9.2 |
| Excalidraw | Benzi Product — 9.8 | CodeGraph — 9.5 | Benzi MCP — 9.4 | Bare CC — 9.1 |
| Tokio | Benzi Product — 9.8 | Benzi MCP — 9.7 | Bare CC — 9.4 | CodeGraph — 9.0 |
| Django | CodeGraph — 9.8 | Benzi Product — 9.7 | Benzi MCP — 9.5 | Bare CC — 9.0 |
| Average | Benzi Product — 9.73 | Benzi MCP — 9.15 | CodeGraph — 9.10 | Bare CC — 8.97 |
Bare CC is Claude Code, Sonnet 5. CodeGraph is bare CC + CodeGraph MCP. Benzi MCP is bare CC + Benzi MCP. Benzi Product is Benzi's own standalone agentic loop (running on Claude Sonnet 5 as well, but no Claude Code). Benchmark methodology replicated from CodeGraph's own published benchmarks: github.com/colbymchenry/codegraph
Our read: we've already shown Benzi is considerably cheaper (2x) and faster (40%) than bare Claude Code on bug-fixing tasks (click here). Code intelligence cutting the cost of localizing and fixing a bug, while raising the cost of an open-ended architectural question, tracks with intuition — the same tools that shortcut a targeted fix invite more exploration on a broad "how does this work" question. Benzi's answers also generally cover more depth, or connect systems that bare Claude Code and Claude Code with CodeGraph as an MCP tool miss.