Skip to content
HN On Hacker News ↗

Benzi — Benchmarks

▲ 10 points 5 comments by tweedler290 1w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

73 %

AI likelihood · overall

AI
14% human-written 86% AI-generated
SEGMENTS · HUMAN 1 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 518
PEAK AI % 83% · §2
Analyzed
Sep 11
backend: pangram/v3.3
Segments scanned
2 windows
avg 259 words each
Distribution
14 / 86%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 518 words · 2 segments analyzed

Human AI-generated
§1 Human · 12%

Benzi · benchmarks Three ways we've measured it. Apples to apples comparison: Benzi harness vs mainstream harnesses, 24 GitHub issues, 10 languages. Benzi harness on SWE-bench Verified. 78.2% of 500 real issues resolved at under 10¢ a fix. Apples to Apples comparison: Code Graph's code intelligence vs Benzi's AI-native code intelligence Every task, every attempt, verbatim — nothing held back. Learn more about Benzi: benzi.fly.dev/about Benzi's KPI (Key Performance Indicator) — source lines read Every harness opens more source as bugs get harder. The question is the slope.

§2 AI · 83%

Each point is one bug; the 24 are laid out easiest to hardest, left to right. Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty. Hover any point for the bug and its count. source lines read · 24 bugs · lowest of the four marked bug BenziSonnet BenziDeepSeek Claude CodeSonnet DeepSeek HarnessDeepSeek mux64643831,372 commons-cli39235456771 addressable1802402001,603 jsoup6117060616 yaml-cpp120383194461 cJSON17187105854 dayjs64336307610 gson7815880516 CsvHelper42390335805 semver3534685631,857 hashie153172300458 money6815726611,892 rich2555147361,421 fmt6043152641,097 sqlglot227115800777 scrapy3861,5801,2592,127 marked9058072,1674,091 http-parser4491,2011,2702,635 zod8411,4761,9563,608 quartznet4881,7071,8974,197 sqlparser6201,0071,2252,334 nlohmann/json3799291,7645,201 ts-pattern1,1301,2321,1391,910 nats-server9892,1492,5832,385 all 249,12516,40720,70443,598 Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The green figure in each row is the lowest of the four. The same 24 bugs in the same order, with wall clock in place of lines read. Wall clock is raw here — unlike the tables above, Benzi's per-repo index build is not subtracted, so these seconds run slightly higher than the warm figures quoted elsewhere on this page. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit. And the same again with dollars on the vertical axis. Priced at the published per-token rates, same run selection as the chart above it. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline — the per-bug figures behind them are in the DeepSeek table further down. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here. cost per fix · USD at list price · lowest of the four marked bug BenziSonnet BenziDeepSeek Claude CodeSonnet DeepSeek HarnessDeepSeek mux$0.39$0.036$0.25$0.023 commons-cli$0.26$0.030$0.29$0.014 addressable$0.57$0.043$0.35$0.042 jsoup$0.23$0.025$0.30$0.051 yaml-cpp$0.25$0.068$0.37$0.024 cJSON$0.27$0.036$0.44$0.073 dayjs$0.21$0.038$0.55$0.040 gson$0.19$0.015$0.45$0.022 CsvHelper$0.25$0.038$0.58$0.042 semver$0.62$0.10$1.43$0.097 hashie$0.69$0.051$0.92$0.043 money$0.79$0.036$0.95$0.052 rich$0.33$0.11$1.30$0.053 fmt$0.54$0.14$1.12$0.071 sqlglot$0.62$0.033$1.23$0.059 scrapy$0.73$0.10$1.35$0.19 marked$1.08$0.20$3.59$0.13 http-parser—$0.36$3.44$0.23 zod$1.53$0.13$2.47$0.33 quartznet$0.95$0.23$3.99$0.31 sqlparser$0.72$0.14$3.22$0.33 nlohmann/json$1.03$0.17$3.68$0.44 ts-pattern$3.74$0.31$4.33$0.053 nats-server$1.99$0.23$2.94— all 24$17.96$2.66$39.54$2.70 Priced at published per-token rates. The green figure in each row is the lowest of the four; the two DeepSeek columns are cheaper largely because that model costs roughly twenty times less per token. Blank cells are the two runs that never produced a fix.