Skip to content
HN On Hacker News ↗

The Complete LLM Leaderboard: The Ed-o-meter

▲ 240 points 124 comments by ed-is-ai 2w ago HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly AI, with some human-written content.

86 %

AI likelihood · overall

AI
7% human-written 93% AI-generated
SEGMENTS · HUMAN 1 of 2
SEGMENTS · AI 0 of 2
WORD COUNT 270
PEAK AI % 60% · §1
Analyzed
Aug 23
backend: pangram/v3.3
Segments scanned
2 windows
avg 135 words each
Distribution
7 / 93%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 270 words · 2 segments analyzed

Human AI-generated
§1 Mixed · 60%

Updated 23 August 2026  ·  Model Evaluation  ·  Ed Yau, Applied AI Architect, Kerv Same driver, same track. The LLM is the star. Seventeen leading models driven round the identical 28-realworld task lap — one harness, same verbatim prompts, deterministic grading — and the results go on the board. Short version: if you run one model, run glm-5.3 — 100% pass, a 9.3 rubric, $0.28 for the lap, about a fifth of gpt-5.5's cost (check with compliance first, though). gpt-5.5 is the faster alternative: 13.2s TTFT versus glm-5.3's 16.3s. gpt-5.6-luna remains the cheapest workhorse for low-risk, retryable jobs; haiku-4-5 if you need it right first time. Choose sonnet-4-6 for quality without the wait.

§2 Human · 21%

The reasoning, with the caveats → 17 models  ·  28 tasks  ·  single trial  ·  latest source run 20260822T172041Z  ·  Change log ← All posts Which Model Tops Our Leaderboard? How the LLMs did in our realworld tests. Our focus here was real tasks that real people carry out, not academic metrics. We focus on single tasks to simplify the assessment. An agentic flow is ultimately a series of such tasks. Think of these like unit tests for the agent. We made them cheap enough to run so that even the whole suite costs just $30. See every task and each model's actual answer, or compare two models head to head → The overall score is the pass rate across my 28 realworld tasks. As we only had a limited number of trials there is a wide Wilson interval — the whiskers on the chart. Summary of results: click a column to sort by your chosen metric.