Skip to content
HN On Hacker News ↗

GitHub - ninjahawk/livenerf: Benchmark for tracking model capability after release.

▲ 923 points • 393 comments • by bryan0 • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly AI, with some human-written content.

96 %

AI likelihood · overall

AI
2% human-written 98% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 47
PEAK AI % 53% · §1
Analyzed
Sep 30
backend: pangram/v3.3
Segments scanned
1 windows
avg 47 words each
Distribution
2 / 98%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 47 words · 1 segments analyzed

Human AI-generated
§1 Mixed · 53%

A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch. 📋 The plan · 📊 Results · 🔬 How it works · 🧪 Pre-registration livenerf is a small, boring, append-only benchmark for one question: does a model get worse after it ships?