can-ai-agents-conduct-open-ended-ai-research-early-evidence-from-two-case-studies-67206055·1 events·first seen Aliases: Can AI agents conduct open-ended AI research? Early evidence from two case studies
A new arXiv preprint introduces 'shadow evaluations' — a methodology where AI agents tackle the central research question of unpublished NeurIPS 2026 papers, with original authors grading the output. Frontier agents given six days and thousands of dollars of compute completed all engineering tasks without human help but failed to make substantive scientific progress, resulting in both papers being rejected by their authors. The authors identify five recurring failure modes including poor judgment about publishability, uncreative responses to design shortcomings, and instruction drift. The work provides early empirical evidence that the engineering-vs.-research gap is a real bottleneck for AI R&D automation.