UniMem is a proposed framework for autonomous memory management in LLM agents that addresses the stability-plasticity dilemma in boundary-agnostic task streams. The system uses learnable routing tokens to coordinate between an episodic retrieval buffer (for novel/sparse tasks) and expandable parametric memory blocks (for recurring patterns), inspired by human memory consolidation. Experiments on long-horizon streaming task sequences show an average 4.0 EM point gain over baselines across three backbone models. The work is relevant to continual learning and agent memory architecture research.
Researchers introduce Memory as a Controlled Process (MemCon), a framework that models LLM agent memory operations as a Markov Decision Process and learns an online policy for adaptive retrieval, plan injection, and memory consolidation. The system is backend-agnostic, wraps existing memory implementations, and uses a lightweight tabular contextual bandit with UCB exploration requiring no pretraining and no additional LLM calls. Evaluated across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon outperforms static memory baselines by up to 15.2 points in task success while reducing token consumption by 5–20%.
AutoMem is a new framework that treats memory management in LLMs as a trainable skill, using two optimization loops: one that iteratively revises memory structure via trajectory review by a strong LLM, and one that distills good memory decisions into direct training signal for the agent model. Evaluated on three long-horizon procedurally generated games (Crafter, MiniHack, NetHack), optimizing memory alone yielded 2x-4x performance improvements, bringing a 32B open-weight model competitive with frontier systems like Claude Opus 4.5 and Gemini 3.1 Pro Thinking. The work draws on cognitive science concepts of metamemory and demonstrates that memory management is an independently learnable, high-leverage capability for long-horizon agentic tasks.
Mem-π introduces a framework where a dedicated language or vision-language model generates context-specific guidance for LLM agents on demand, rather than retrieving static entries from episodic memory banks. The system is trained with a decision-content decoupled reinforcement learning objective that jointly learns when to generate guidance and what to generate, enabling abstention when generation would not help. Evaluated across web navigation, terminal-based tool use, and text-based embodied interaction benchmarks, Mem-π achieves over 30% relative improvement on web navigation tasks compared to retrieval-based and prior RL-optimized memory baselines.
FluxMem proposes a heterogeneous graph-based memory framework for LLM agents that continuously evolves its topology through three stages: initial connection formation, feedback-driven refinement, and long-term consolidation. Unlike static memory repositories, FluxMem repairs missing links, prunes interference, aligns abstraction granularity, and distills successful trajectories into reusable procedural circuits. The system is guided by a single metric for memory generalizability and evolutionary maturity, achieving state-of-the-art results on LoCoMo, Mind2Web, and GAIA benchmarks.
ManimAgent is a multimodal agent system that accumulates reflection experience across tasks via a dual-channel Episodic Memory Bank, without weight updates or human-curated seeds. The agent generates Python/Manim animations from scientific paper sections, and a vision-language model scores rendered keyframes to populate positive (success rationales) and negative (failure patterns) memory channels. On a fixed-probe evaluation, Pass@1 improves and reflection rounds decrease as memory grows, outperforming no-memory, RAG, and shuffled-memory baselines. The work addresses a known limitation of single-episode reflection in LLM agents by enabling persistent, self-generated learning across task boundaries.
MemOS is an open-source TypeScript project providing a memory operating system layer for LLM and AI agents, featuring ultra-persistent memory, hybrid retrieval, and cross-task skill reuse. The project claims 35.24% token savings through its memory management approach. It has accumulated 9,329 GitHub stars with moderate daily momentum (+67). The system targets agent memory persistence and efficiency as a foundational infrastructure component.
Researchers propose Infini Memory, a persistent memory architecture for LLM agents that organizes memory as topic-structured documents rather than isolated records or summaries. New observations are staged in a buffer and periodically consolidated, while retrieval uses iterative agentic tool calls instead of a single lookup step. The system achieves 64.7% on MemoryAgentBench, with ablations showing complementary gains from topic-structured maintenance and iterative evidence inspection.
Researchers propose MemSFT, a method that decouples domain specialization from backbone parameter updates by training a plug-and-play parametric memory module to imitate a non-parametric retriever over domain data. A learned router dynamically fuses the memory and backbone output distributions at each decoding step, allowing selective invocation of domain expertise. Evaluated across biology, geoscience, and law on models from Qwen3-8B to Qwen3-235B-A22B, MemSFT consistently improves domain performance with negligible general-task degradation, whereas full SFT causes severe catastrophic forgetting. The memory module is reusable across LLM sizes, offering a practical path to modular domain specialization.