Skip to content
HN On Hacker News ↗

ProgramBench Vetted

▲ 27 points 1 comments by rigelbm 3d ago HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

79 %

AI likelihood · overall

AI
17% human-written 83% AI-generated
SEGMENTS · HUMAN 0 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 128
PEAK AI % 79% · §1
Analyzed
Aug 22
backend: pangram/v3.3
Segments scanned
2 windows
avg 64 words each
Distribution
17 / 83%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 128 words · 2 segments analyzed

Human AI-generated
§1 AI · 79%

Part of the Computer Anthology benchmark family. ProgramBench Vetted Can an agent rebuild a program from a runnable binary? That is the fundamental question behind ProgramBench. We are releasing ProgramBench Vetted, a set of 50 tasks built around the same premise, with improvements to the task design process that make the benchmark fairer and more reliable. In addition to preserving measurable partial progress and deterministic test based grading, our task design process adds a broader set of controls for failure modes involving test duplication, environment quality, hackability, and fairness.

§2 Mixed · 60%

ProgramBench Vetted model results Tasks passing 100% of tests, pass@1 across 50 held out tasks · default ProgramBench mini-swe-agent configuration ProgramBench matters because program reconstruction is one of the relatively few coding benchmark formats built around very long-horizon work.