Skip to content
HN On Hacker News ↗

ORCA-bench: How Ready Are Language Model Agents for Oncall?

▲ 30 points 11 comments by yruzin 3w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

59 %

AI likelihood · overall

Mixed
21% human-written 79% AI-generated
SEGMENTS · HUMAN 1 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 283
PEAK AI % 73% · §1
Analyzed
Jul 31
backend: pangram/v3.3
Segments scanned
2 windows
avg 142 words each
Distribution
21 / 79%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 283 words · 2 segments analyzed

Human AI-generated
§1 AI · 73%

View PDF HTML (experimental) Abstract:Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and full source-code access--with 1,079 RCA tasks that systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $\kappa_w=0.90$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard--a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric. Crucially, these are performances on a curated 50 GB / six-day testbed with tasks investigated in isolation on a system whose code and instrumentation are public. Since real production systems are order of magnitudes larger, more dynamic, and more idiosyncratic, the gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability.

§2 Human · 5%

We release the public set at this https URL. Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE) Cite as: arXiv:2607.28545 [cs.CL] (or arXiv:2607.28545v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2607.28545 arXiv-issued DOI via DataCite (pending registration) Submission history From: Albert Gong [view email] [v1] Thu, 30 Jul 2026 17:14:07 UTC (220 KB)