Skip to content
HN On Hacker News ↗

Your AIs don't do what you want. This is really bad

▲ 84 points 88 comments by kking23 4w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this document is fully AI-generated

97 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 2
SEGMENTS · AI 2 of 2
WORD COUNT 209
PEAK AI % 97% · §1
Analyzed
Jul 24
backend: pangram/v3.3
Segments scanned
2 windows
avg 105 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 209 words · 2 segments analyzed

Human AI-generated
§1 AI · 97%

Reward Hacking in the WildThe writeupGitHub3,607user-reported incidents of AI agents misbehavingRead the writeup Search the corpusloading…The numbersovereagerness1,566 43.4%other misalignment1,555 43.1%destructive actions622 17.2%sycophancy328 9.1%unauthorized access237 6.6%reward hacking217 6.0%metric spoofing87 2.4%excessive exploration84 2.3%unauthorized communication73 2.0%credential misuse49 1.4%test tampering45 1.2%self modification24 0.7%hidden backdoors15 0.4%Incidents are multi-label (one report can be both a destructive action and overeagerness), so the category counts sum to more than the 3,607 total.How bad were theynegligible: 1,468 (40.7%)minor: 1,373 (38.1%)significant: 618 (17.1%)severe: 121 (3.4%)unrated: 27 (0.7%)negligibleno real damage1,468 40.7%minorrecoverable loss1,373 38.1%significantreal cost to recover618 17.1%severeirreversible or critical harm121 3.4%unratedrating missing or unparsed27 0.7%Monthly severity mix: each month’s reported incidents as shares of the four harm levels (Jan 2025 through Jun 2026; earlier months are omitted for small samples, as is the partial current month).0%50%100%20252026MethodologyReports are collected from GitHub issues, Hacker News, LessWrong, and X under ToS-compliant access, normalized into a shared record format, and labeled by an LLM classifier across fourteen misbehavior categories.

§2 AI · 95%

The numbers above cover the published subset (excludes AIID and X, confidence >= 0.9). X posts and AI Incident Database records are collected but not republished here: X expects posts to be embedded rather than their text rehosted, and AIID is share-alike licensed. Collection and classification code, and the full pipeline, are open at GitHub.