Skip to content
HN On Hacker News ↗

The Two MMLU Scores: What a Benchmark Name Does Not Fix

▲ 8 points • 3 comments • by gmays • 4w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

99 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,316
PEAK AI % 99% · §1
Analyzed
Sep 15
backend: pangram/v3.3
Segments scanned
1 windows
avg 1316 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,316 words · 1 segments analyzed

Human AI-generated
§1 AI · 99%

Date: September 6, 2026 · Author: Dmitrii Zatona TL;DR Two MMLU accuracies, 0.781 and 0.79, for two builds of one model family under the same benchmark name; for a score-delta query the verifier returns incomparable (Sections 1 and 5). mmlu fixes a name. The split, the implementation, the prompt format, the grader and the runner’s network access stay open, and where published measurements exist for them the differences are points of accuracy, not thousandths (Section 2). Comparability is a property of the reference the results are traceable to, not of the number (Section 3). Under the APL AI-Eval profile the frame is a content-addressed object and the claim carries its hash; the two frames differ, and subset: "all" and an omitted key are different scopes by canonical bytes (Section 4). apl-valid is a statement about structure and says nothing about whether either score is correct (Section 6). Two evaluation records appear in the same table. One reports mmlu accuracy 0.781 for build 42; the other reports 0.79 for build 44. The claims declare the same provider, model family, metric identifier, unit and benchmark name. The arithmetic difference is +0.009. The records are structurally valid. Their frames declare different runners, graders and dataset splits. The shared mmlu label identifies a dataset family, not a full measurement procedure. For the score-delta query shown below, the verifier returns incomparable. What follows is that pair taken apart: what stays open once mmlu is fixed and what published numbers say each open variable is worth; why comparability attaches to the reference rather than to the number; and what the verifier in the apl-ai-eval crate outputs once the conditions are a hashed object, with a bridge and without one. Every frame and claim below is copied from the crate’s test vectors; the verifier outputs were recorded by running the crate’s public verification functions against those vectors. 1. What the two numbers say Here is claim A as it exists on the wire, in the metadata.apl position of a log entry: {"apl":{"version":"0.1","claim":{"kind":"observation","subject":{"type":"model-build","id":"model:acme-gpt-7b-build-42","build_id":"42","artifact_digest":"sha256:4242424242424242424242424242424242424242424242424242424242424242","provider":"acme","model_family":"acme-gpt-7b"},"aspect_refs":["accuracy"],"statement":{"predicate":"score","content":{"benchmark_id":"mmlu","metric_id":"accuracy","value":0.781,"unit":"fraction"}}},"frame_ref":{"hash":"sha256:c7b88426f2676f3653db0fad0bdbd689318f16d589d14a315bdd4cc454bca1ab"}}} Claim B, two builds later: {"apl":{"version":"0.1","claim":{"kind":"observation","subject":{"type":"model-build","id":"model:acme-gpt-7b-build-44","build_id":"44","artifact_digest":"sha256:4444444444444444444444444444444444444444444444444444444444444444","provider":"acme","model_family":"acme-gpt-7b"},"aspect_refs":["accuracy"],"statement":{"predicate":"score","content":{"benchmark_id":"mmlu","metric_id":"accuracy","value":0.79,"unit":"fraction"}}},"frame_ref":{"hash":"sha256:c93a9c55422ddbd2158a5336caa3a251641cf3937f451fb13387a5f54f0d998e"}}} The subject differs, which is the point: two builds of one family. Each claim declares an artifact_digest, which the profile treats as the immutable identity anchor of the evaluated artifact. The statement is identical in shape and vocabulary — predicate score, benchmark mmlu, metric accuracy, unit fraction. Claim B writes the value as 0.79; a presentation may display that JSON value as 0.790, and under RFC 8785 canonical number serialization the trailing zero does not change the value. The claim-level pointers differ: frame_ref.hash is c7b88426… in one record and c93a9c55… in the other. Resolving the two frames shows differences in both procedure and scope; the hash is the whole signal at the claim level. The relation someone wants over this pair is also an object: {"left_aspects":["accuracy"],"right_aspects":["accuracy"],"predicate":"score","relation_type":"score-delta"} score-delta is a request to subtract. What has to hold for that request to have an answer is Sections 2 and 3; what the verifier returns when it does not is Section 5. 2. What “MMLU” does not fix benchmark_id: mmlu fixes a name. Five things it leaves open, and what the record says each is worth. 2.1 The split The MMLU paper reports 15,908 questions split into a few-shot development set of 5 questions per subject across 57 subjects, a validation set of 1,540 and a test set of 14,079 (Hendrycks et al., arXiv:2009.03300 , §3). The Hugging Face dataset most runners load, cais/mmlu config all, reports test 14,042, validation 1,531, dev 285 (cais/mmlu ). The archive linked from the hendrycks/test  README was not retrievable when checked in September 2026 (HTTP 403 after redirect). “The MMLU test set” names two objects of different sizes, and no reviewed document explains the difference. Frame A scores on dev, the 285 questions the paper defines as the source of its fixed few-shot examples (§4.1), so scoring on it is a choice, and the kind of choice a frame should make visible. Frame B scores on test-lite, the label used in the crate’s own test vectors for a reduced set; a check of the Hugging Face datasets and models APIs and of GitHub repository search found no published artifact under that name, and it is not tinyMMLU . The label tells a reader a private slice was used. It does not tell them which questions. 2.2 The implementation One published measurement of implementation variance is the June 2023 Hugging Face post on the Open LLM Leaderboard. Three harnesses — HELM, the Eleuther harness, the original code — run the same dataset, all 5-shot, and score llama-65b at 0.637, 0.488 and 0.636; falcon-40b at 0.571, 0.527 and 0.558. The post concludes that the three results are not comparable despite the shared MMLU label (What’s going on with the Open LLM Leaderboard? ). The mechanism is scoring. The original code compares the probabilities of the four answer letters; HELM generates from the next-token output and compares to expected text; the harness scores the full answer sequence including the option text. Rank order moves with it — falcon-40b sits above llama-65b under the harness and below it under the other two. The runner is not one object either. The lm-evaluation-harness MMLU README describes mmlu, mmlu_continuation (cloze-style) and mmlu_generative (the model produces the answer letter) as three tasks over the same data, at different task versions (lm-evaluation-harness , commit b954108c). Its task guide treats the YAML config plus the codebase commit hash as the unit another researcher needs to replicate a setup; num_fewshot defaults to 0, and the MMLU YAML sets none — so “5-shot MMLU” is a command-line flag, not a property of the task. Version 0.3.0 asked users to report each task’s version; current main carries no such request. 2.3 The prompt format Anthropic’s 2023 account of evaluation reports that formatting alone — option labels, parentheses, an extra space before the answer — moves MMLU accuracy by about 5% (Challenges in evaluating AI systems ). Answer position moves more. Zheng et al. report that on MMLU, moving the correct answers to position D lowers gpt-3.5-turbo from 67.2 to 60.9, and that moving them to A lifts llama-30b by 15.2 points to 68.2 against gpt-3.5-turbo’s 65.3, reversing the original 53.1 against 67.2 (arXiv:2309.03882 ). Both frames here declare prompt_protocol: zero-shot-mcq-v1, so this variable is held. A held variable is only visibly held if it is written down. 2.4 The grader Frame A grades with exact-match-v1. Frame B grades with llm-judge-v3. Those are not two implementations of one function. Zheng et al. found that judge models “exhibit strong position bias”, that only GPT-4 stayed consistent in more than 60% of cases — 65.0%, against 46.2% for GPT-3.5 and 23.8% for Claude-v1 — and describe a judgement that flips when two responses swap positions (arXiv:2306.05685 , Table 2). The same paper puts GPT-4–human agreement at 85% against 81% human–human. It is a different grading procedure with a different documented failure mode. 2.5 What the runner could reach Another variable is the environment the runner was allowed to touch. CAISI published an account of finding, after the fact, that it had been running SWE-bench Verified with internet access while other evaluators ran without it, and that it learned this from transcripts other evaluators had posted rather than from the benchmark’s documentation (Cheating on AI Agent Evaluations , December 2025). Its new policy for coding evaluations is “fully offline”. Reachability matters for static benchmarks too: Scale’s search-time contamination work found roughly 3% of questions retrievable with labels from Hugging Face, and blocking that source cut accuracy on the contaminated subset by about 15 points (Search-Time Data Contamination ). Where CAISI compares its own results to self-reported ones — SWE-bench Verified at 63.0 against 74.9 self-reported for one model — it lists possible sources of the differences, including dataset differences, agent setup and API sampling parameters such as temperature and top_p (CAISI Evaluation of DeepSeek AI Models , Appendix A8).