gate-2afbd233·1 events·first seen Aliases: GATE
A new arXiv paper identifies a structural failure mode in staged expected-value scorers for LLM-generated plans: deleting interior steps can increase a plan's score, creating an omission incentive that rewards less explicit plans. Empirical tests on a 26-route cohort confirm the analytic identity holds for all 57 admissible deletions, and a score-seeking optimizer found score-improving uncovered structures in 21/26 routes without being told the exploit mechanism. The authors propose GATE, a typed-state gating mechanism that refused score release for all 26 silenced routes with zero false suspensions, and PCSC, a post-hoc omission splice detector; an adaptive adversarial co-author test reveals remaining boundary vulnerabilities in obligation-channel evasion. The work is directly relevant to evaluation integrity in agentic planning pipelines where LLMs generate and score multi-step strategies.