Skip to content
HN On Hacker News ↗

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

▲ 104 points 131 comments by doppp 3w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

2 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 238
PEAK AI % 2% · §1
Analyzed
Aug 4
backend: pangram/v3.3
Segments scanned
1 windows
avg 238 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 238 words · 1 segments analyzed

Human AI-generated
§1 Human · 2%

Authors:Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman View PDF HTML (experimental) Abstract:Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of the our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches. Comments: Accepted at ICML 2026 Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2602.16763 [cs.AI] (or arXiv:2602.16763v3 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2602.16763 arXiv-issued DOI via DataCite Submission history From: Mubashara Akhtar [view email] [v1] Wed, 18 Feb 2026 16:51:37 UTC (222 KB) [v2] Sat, 30 May 2026 16:41:50 UTC (640 KB) [v3] Mon, 29 Jun 2026 17:01:58 UTC (636 KB)