Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
AIArticle text · 128 words · 2 segments analyzed
Part of the Computer Anthology benchmark family. ProgramBench Vetted Can an agent rebuild a program from a runnable binary? That is the fundamental question behind ProgramBench. We are releasing ProgramBench Vetted, a set of 50 tasks built around the same premise, with improvements to the task design process that make the benchmark fairer and more reliable. In addition to preserving measurable partial progress and deterministic test based grading, our task design process adds a broader set of controls for failure modes involving test duplication, environment quality, hackability, and fairness.
ProgramBench Vetted model results Tasks passing 100% of tests, pass@1 across 50 held out tasks · default ProgramBench mini-swe-agent configuration ProgramBench matters because program reconstruction is one of the relatively few coding benchmark formats built around very long-horizon work.