grafana-bdfb3382·1 events·first seen Aliases: Grafana
ORCA-bench is a new benchmark that tests general-purpose coding agents on oncall root cause analysis (RCA) in a production-fidelity environment, using a live OpenTelemetry-instrumented microservice system with 1,079 tasks spanning metrics, logs, traces, and source code. Across five frontier agents, the best RCA accuracy is 25.3% on medium-difficulty tasks and 10.0% on hard tasks, with even Claude Fable 5 failing to close the gap significantly. The benchmark is notable for its expert-curated ground truth (Cohen's κ_w=0.90 human-LLM agreement) and its argument that reported performance is a lower bound on real-world difficulty, since production systems are far larger and more dynamic than the testbed.