A new arXiv paper introduces the concept of the 'regression tax' — the phenomenon where adding procedural skills to LLM agents causes previously-solved tasks to fail, even when the skill is never invoked. Across nearly 6,000 runs on two office automation benchmarks and three model harness stacks, the authors identify three regression mechanisms: skill description osmosis (context pollution), grounding displacement (procedure overrides input interpretation), and verification displacement (procedure suppresses output checking). The key finding is that top-performing skill configurations succeed primarily by regressing less, not by gaining more, and that grounding and verification are more important to reliability than procedural skill choice.
A new arXiv paper analyzes how biased LLM judges used as reward signals in self-evolving agent systems can silently disable the 'skill retirement' mechanism that prevents skill libraries from degrading below a no-skill baseline. The authors show that false-pass bias—where failures are incorrectly scored as passes—disables contribution-based retirement past a sharp threshold that cannot be overcome with more data, while leaving aggregate performance metrics unchanged, making the failure invisible. The paper proposes a defect-injection audit to determine pre-deployment whether a judge falls above or below the critical threshold.
A new arXiv preprint tests the implicit assumption that LLM evaluation is easier than generation, using a controlled in-context QA setup across four benchmarks (SQuAD 2.0, DROP, HotpotQA, MuSiQue) and two models. Results show generation accuracy exceeds self-evaluation accuracy on three of four benchmarks, with attention analysis revealing that evaluation attends to context 3–5x less than generation does. LoRA fine-tuning experiments confirm the asymmetry is not a training artifact, with cross-task interference observed in both directions. The findings directly challenge assumptions underlying LLM-as-a-Judge and self-evaluation pipelines widely used in RLHF and agentic systems.
A new arXiv preprint introduces Double Ratchet, a system that co-evolves both evaluation metrics and agent skills in settings where no reliable automatic verifier exists. The metric loop uses evolutionary search over small drawback detectors anchored to a small reference set, while the skill loop uses a lifecycle-managed approach; together they retain 88–110% of the performance lift achievable with ground-truth metrics across code generation (MBPP+), text-to-SQL (Spider 2.0-Snow), and report generation tasks. The paper also addresses safety, showing that anchor discipline and outer audits can catch and repair cases where evolved skills game the rubric. This work directly addresses a core bottleneck in self-improving agent systems: the chicken-and-egg problem of needing a reliable evaluator to improve.
SkillGenBench is a new benchmark designed to evaluate the ability of LLM agents to generate correct, reusable, and executable skills from raw repositories and documents, rather than merely using pre-provided skills. It covers two generation regimes (task-conditioned and task-agnostic) and two procedural sources (repository-grounded and document-grounded), with standardized execution-based evaluation protocols. Experiments across multiple skill-generation methods reveal substantial performance variation and distinct failure modes depending on source type. The benchmark aims to establish skill generation as an independent research problem within agent systems.
Researchers introduce OpenSkillRisk, a safety benchmark of 263 risky third-party skills drawn from public skill marketplaces, designed to evaluate how well LLM-based agent systems recognize and avoid latent execution-time risks. Experiments across three CLI agent frameworks and thirteen LLMs show no system handles risky skills reliably, with even the safest configurations executing unsafe actions in ~17% of cases. The benchmark identifies three recurring failure patterns: failure to recognize risk, recognizing risk but acting anyway, and over-following skill instructions beyond user intent. Context-dependent and system-level risks prove especially difficult for current agents to handle.
A new arXiv paper argues that standard LLM benchmarks overstate model capabilities by focusing on average performance on training-data-adjacent tasks while ignoring response variance and error magnitude. The authors introduce a novel benchmark requiring frontier LLMs to write code for data analysis tasks, comparing results against human expert submissions. Human experts outperformed the frontier LLM on average across multiple metrics and showed lower performance variability. The findings challenge the prevailing narrative that LLMs perform at human-expert level on knowledge economy tasks.
A new arXiv preprint introduces SkillComposer, a method that frames skill selection for LLM agents as a structured prediction problem — jointly deciding which skills to activate, how many, and in what order via a constrained autoregressive decoder over skill identifiers. The approach addresses a bottleneck in growing skill libraries where existing retrieval and full-context methods fail to capture the joint nature of skill composition. Evaluated on SkillsBench across two production-grade coding agents (GPT-5.2-Codex and Gemini-3-Pro-Preview), SkillComposer raises pass rates by +23.1 and +18.2 percentage points over no-skill baselines, matching gold-skill retrieval upper bounds at lower prompt-token cost.
A new arXiv paper systematically evaluates a range of LLM conditioning methods across both concept injection and removal scenarios, finding that efficient steering methods often degrade fluency significantly. A key finding is that activation steering is substantially less effective on instruction-tuned models than on base models, a previously overlooked interaction. Simple prompting and supervised fine-tuning work for concept injection but not removal, and cheap textual metrics are found to correlate well with expensive LLM-as-judge evaluations.