desktop-delta-bench-37a9b813·1 events·first seen Aliases: Desktop-Delta Bench
Researchers introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances designed to test whether computer-use agents can correctly interpret causal GUI transitions produced by actions — a capability distinct from end-task success or single-frame grounding. The benchmark covers ~15 Linux applications across 50 task domains and targets three failure dimensions: state verification, source tracking, and context-aware control. Evaluation of 8 closed and open-source model families reveals consistent gaps, with best temporal-ordering exact-match rates at ~65% and action-family inference notably harder than action localization. DDB fills a diagnostic gap between GUI grounding benchmarks and end-to-end task success metrics.