memprobe-d8c5a3a4·2 events·first seen Aliases: MemProbe
MEMPROBE is a new benchmark that evaluates long-term memory in LLM agents by treating memory as an auditable artifact rather than measuring it only through downstream task performance. After a memory-equipped agent assists simulated users across a trajectory of tasks, the benchmark attempts to reconstruct a hidden, taxonomy-anchored user-state bank from the agent's memory store. Testing across 5 memory systems and 50 simulated users with 31 hidden dimensions each, the authors find that task completion and memory recovery are largely independent capabilities — task success nearly saturates even for memoryless baselines, while structured user-state recovery remains moderate (~0.6) and degrades under top-k retrieval constraints.
AgentCL is a new benchmark and evaluation framework designed to rigorously assess continual learning in language agents, addressing gaps in existing benchmarks that focus on retrieval over long-context documents or use naive task streams with limited cross-task analysis. The framework constructs compositional task streams where earlier sub-solutions, evidence, or workflows are intentionally reusable in later tasks, contrasting them with naive streams to measure transfer gains. The authors also introduce MemProbe, a probing method that stores interactions, insights, and skills while filtering unreliable experiences during consolidation. Empirical results across coding, deep research, and language understanding tasks show that controlled streams better distinguish memory design quality, and that naive streams can mask memory-induced degradation.