inmind-6b2e1289·1 events·first seen Aliases: InMind
Researchers introduce InMind, a 125-task expert-verified benchmark targeting a failure mode in agent long-term memory: facts stored in memory are not retrieved when the query lacks surface-level textual overlap with the relevant memory (e.g., a nut allergy should affect a macaron recommendation but shares no keywords). Testing six vector, graph, and agentic memory systems reveals a stark gap — backbone models answer 84% of indirect queries when the relevant memory is placed in context, but retrieval systems surface the correct answer at most 14.4% of the time despite near-perfect on-demand recall. The study isolates the failure to the query-conditioned retrieval interface itself, framing routing — deciding which facts must remain visible — as the core open problem.