Entity · other

language agents

otheractivelanguage-agents-0d18426e·2 events·first seen May 25, 2026

Aliases: language agents

Co-occurring entities

AgentCL MemProbe Continual Learning non-parametric memory Model-Generated Agent Skills (paper)skill extraction meta-skill negative transfer

More like this (12)

hello-agents Semantic Agent TradingAgents AI Agents tool-augmented language agents video agents computer-use agents LLM agents GUI Agents coding agents tool-calling agents ai-agent-book

Recent events (2)

6arXiv · cs.CL·Jun 2, 2026·source ↗

AgentCL: A Rigorous Evaluation Framework for Continual Learning in Language Agents

AgentCL is a new benchmark and evaluation framework designed to rigorously assess continual learning in language agents, addressing gaps in existing benchmarks that focus on retrieval over long-context documents or use naive task streams with limited cross-task analysis. The framework constructs compositional task streams where earlier sub-solutions, evidence, or workflows are intentionally reusable in later tasks, contrasting them with naive streams to measure transfer gains. The authors also introduce MemProbe, a probing method that stores interactions, insights, and skills while filtering unreliable experiences during consolidation. Empirical results across coding, deep research, and language understanding tasks show that controlled streams better distinguish memory design quality, and that naive streams can mask memory-induced degradation.

Long Context Evolution Evaluation and Benchmarking AgentCL MemProbe Continual Learning +3 more

6arXiv · cs.AI·May 25, 2026·source ↗

Systematic Study of Model-Generated Agent Skills Across the Full Skill Lifecycle

This paper presents a utility-grounded evaluation framework for model-generated agent skills, covering the full lifecycle of experience generation, skill extraction, and skill consumption across five agentic task domains. The authors find that while such skills are beneficial on average, they exhibit non-trivial negative transfer, and that skill utility is independent of model scale or baseline task strength. A key finding is that strong extractors are not necessarily strong consumers and vice versa. The work culminates in a 'meta-skill' that guides extraction toward utility-correlated features, consistently improving skill quality and reducing negative transfer.

Evaluation and Benchmarking Agent and Tool Ecosystem Model-Generated Agent Skills (paper)skill extraction meta-skill +2 more