Skip to content
HN On Hacker News ↗

AGI Ranker - Open AGI Score for Frontier AI Models

▲ 4 points 3 comments by baraklaniado 4w ago HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly AI, with some human-written content.

97 %

AI likelihood · overall

AI
1% human-written 99% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,733
PEAK AI % 97% · §1
Analyzed
Jul 30
backend: pangram/v3.3
Segments scanned
1 windows
avg 1733 words each
Distribution
1 / 99%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,733 words · 1 segments analyzed

Human AI-generated
§1 AI · 97%

v2.0.6 · 2026-07-30 · Our social preview card was still showing pre-v2 numbers, weeks after the numbers changed. Anyone sharing a link to this site saw a card reading 14 live benchmarks, 20 models tracked and “multimodal”. The truth is 10, 21 and Visual Reasoning. The image file on our server was in fact correct and had been regenerated with the rest of the v2 release; what was wrong was our assumption that replacing a file replaces what the world sees. Social platforms cache preview images against the URL, so re-rendering the same filename changes nothing for anyone who has already shared the link, and we had verified the file rather than the card. Every preview image now lives at a versioned URL, the homepage card and all 21 model cards, which forces every platform to fetch it again. The previous files stay where they are, so an existing embed shows an old image rather than a broken one. The card generator carries the version token now, so this is a one-line change next time rather than a rediscovery. This is the same failure as the human-ceiling counts two releases ago: a rendered artifact that nobody re-derived after the thing underneath it changed. The site text was clean, and we checked: no page description, title or preview text anywhere on the site still carries the old figures. v2.0.5 · 2026-07-29 · Tapping a specialty tab on a phone gave you a wall of text instead of scores. The specialty-index notice rendered 586px tall on a 375px screen, 72% of the viewport, and pushed the leaderboard 742px below the tab bar. A reader tapping Coding to see coding scores got a screenful of methodology and had to scroll to find a single number. Below tablet width these notices now collapse to a summary with a More expander, and the full text is one tap away in place. Collapsed, the notice is 151px and the table sits 307px from the tabs, so it is reachable in one short scroll. Nothing was deleted and nothing was softened. The visible summary is a compression, not a gentler version: it still says the specialty index is a 0 to 10 scale rather than the AGI Score, that 10 means a perfect score on every benchmark in the set, and that this is not a human-parity claim. Those are the three things a reader has to know before reading the number, so they stay visible whether or not anybody taps More. The same treatment went to the Value view note about cost being an estimate rather than a measured bill, which kept its estimate caveat in the collapsed line. Switching tabs always re-collapses, so no tab inherits the previous one's expanded state. The expander is a real button with aria-expanded, works without hover, and the Back to AGI Score control stays visible while collapsed. Desktop is unchanged: it has the room, so it shows the full text and no expander at all. v2.0.4 · 2026-07-29 · The mobile tap states would not have rendered on an iPhone. The new section pills carry a pressed state, and it verified correctly in desktop Chromium, which is exactly the wrong place to check it. iOS Safari declines to apply the CSS active state to a link unless a touch listener exists on the element or one of its ancestors, so on the only kind of device that has a touch screen the pills would have looked inert when tapped. One empty listener on the page body fixes it for every tap target on the site. Recorded because verifying a touch behaviour in a mouse browser and calling it done is the kind of shortcut that ships a defect, and because the same trap applies to anything interactive we add from here. v2.0.3 · 2026-07-29 · The site had no mobile navigation at all. On a phone the header rendered a wordmark and a Contribute button and nothing else. The six section links were desktop-only and the version pill was hidden below tablet width, so a visitor on a phone had no route to Corrections, Methodology or anything else, on a page that is otherwise one continuous scroll, and could not see which release they were reading. Small screens now get a compact sticky bar with the version pill visible and a scrolling row of section pills beneath it, carrying the same six anchors as the desktop nav. A row rather than a hamburger: there is no open state to get stuck, nothing to trap keyboard focus, and every section stays one tap away instead of two. Phones also get a back-to-top button once you are far enough down that scrolling back is a chore. Section anchors were landing behind the header at every width, because nothing on the page set a scroll margin; jumping to Corrections put its own heading underneath the sticky bar. Fixed for both layouts. Nothing else moved: the desktop bar is unchanged, and no score, weight or benchmark is touched by this release. v2.0.2 · 2026-07-29 · We were overstating how much of the scale is anchored to humans. Three places on this site claimed that three or four scored benchmarks carry a measured human ceiling, and named OSWorld, FrontierMath and SimpleBench among them. The correct number is one. Of the ten benchmarks on the board, only GPQA Diamond (0.81) carries a measured human ceiling; the other nine are scored against the benchmark maximum, where 100 means a perfect score and no human claim is made at all. The other three did pass the same ceiling audit, and none of them is on the board: OSWorld (0.72) was retired in this very release, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. The claim was wrong in the AGI definition modal, in the methodology summary and in the full methodology, and the three did not even agree with one another. This is the same class of error as the correction we published two versions ago: a number that survived because nobody re-derived it after the thing underneath it changed. The AGI Score is a mixed scale, and it is a good deal more mixed than we were saying. Also in this release: the DeepSeek apology is signed by Barak Laniado, founder and CEO, and its corrections contact is now an email address you can actually write to rather than a link back into the site. The calibration constants box no longer reads “Current as of v1.4” under a v2 banner; the constants are unchanged and were re-confirmed for v2.0.0, which is what it now says. Three hover styles on the corrections card and twelve layout classes on the apology page were also missing from the compiled stylesheet, which is purged to what the leaderboard uses, so they had been failing silently. v2.0.1 · 2026-07-29 · The DeepSeek apology gets its own page. The correction we published in v2.0.0, an ARC-AGI-2 score attributed to DeepSeek V4 Pro that carried no source at all, now has a dedicated page at /corrections/deepseek-arc-agi-2, linked from its entry in the log above. An apology buried as one card among several is easy to walk past, and this one should be readable, citable and linkable on its own. The page sets out what we published, why it was wrong, what it did and did not affect, and what changed in the process so that it cannot recur: under methodology v2 a recorded source is a condition of scoring rather than an expectation, and all 156 scored cells carry one. The wording of the apology itself is unchanged and identical in both places. One count in that entry was also wrong. It said three further cells were voided in the same pass and then referred to four of them in the next sentence. Four further cells were voided, five in total including the DeepSeek one. Corrected here rather than quietly. v2.0.0 · 2026-07-29 · Methodology v2. Every score falls, and no model got worse. Two changes drive it. First, human ceilings. Because a score is normalised as raw divided by a human ceiling, that ceiling is not only an anchor for what counts as human level, it is a volume knob: it multiplies both the level and the spread of a benchmark, and so how hard that benchmark pushes on its component. Most of ours had no measurement behind them. Humanity’s Last Exam sat at 0.50, a number our own internal audit described as a policy floor rather than a finding, and that setting quietly made HLE count double. From now on a ceiling is either a published human result under a protocol comparable to the models’, or it is simply the benchmark maximum and we make no human claim at all. Only four survived the test: GPQA Diamond at 0.81, OSWorld at 0.72 (which we raised, having carried 0.85 against our own recorded measurement), FrontierMath at 0.35, and SimpleBench at 0.837, where we had rounded a human baseline up. Eleven moved to 1.00. Three component scores that previously sat above 100, on a scale whose 100 is meant to be the human mark, no longer do. Second, Agency was rebuilt on four legs across three independent evaluators: SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking and LiveBench Agentic Coding. Terminal-Bench now comes from vals.ai rather than Artificial Analysis, cutting our reliance on any single evaluator. τ³-Banking enters at benchmark maximum, so its true range shows: the best model completes about a third of these stateful banking workflows. Retired: SWE-bench Pro, OSWorld, BrowseComp, Tau-bench retail and airline, Aider Polyglot, Terminal-Bench 2.0. MCP Atlas demoted to informational. The Tool Use tab is withdrawn until a second clean non-coding benchmark exists. Scores fell by about 6 points from the ceiling change and further from the Agency rebuild; ordering shifted only locally, and the top two are unchanged. Third, evaluator concentration. Artificial Analysis had been supplying 45% of the whole board, 97% of Knowledge and 100% of Visual Reasoning, so one evaluator changing terms would have taken two components to zero. GPQA Diamond and MMMU-Pro moved to vals.ai alongside Terminal-Bench 2.1, taking Artificial Analysis to 26% of the board and splitting Knowledge 52/48 between two evaluators. Every switch was measured before it was made rather than after, and each offset is published rather than absorbed: GPQA Diamond −0.30pp with 1.38pp scatter across all 20 models, MMMU-Pro +5.71pp with 1.67pp