workspace-bench-lite-30c5a8b1·1 events·first seen Aliases: Workspace-Bench-Lite
WorkSurface-Bench is a new benchmark of 1,151 tasks evaluating whether enterprise agents correctly select among heterogeneous knowledge sources (documents, tables, dependency graphs) before answering—a capability the authors term 'surface routing.' Evaluating four model backbones across six agent settings, the benchmark reveals a critical gap: agents achieve near-perfect route selection (98.7–99.8% Route F1) under gold-constrained tool access but only 56.1–75.3% answer accuracy, demonstrating that correct surface selection is necessary but insufficient for task completion. The dataset, pipeline, scoring code, and agent harness are released publicly.