Explainers
Unread — 3
paper · Nov 2025
Evo-Memory: Benchmarking Agent Test-time Learning
Separates conversational recall from experience reuse, then measures whether agents actually improve across a task stream.
arXiv 2511.20857 paper · Jul 2025MemoryAgentBench: Evaluating Memory in LLM Agents
Breaks memory into four testable competencies and shows every agent type still fails multi-hop conflict resolution.
arXiv 2507.05257 code · v0.16.8, May 2026Letta — platform for stateful agents
The MemGPT lineage productionised — a working reference for memory blocks and self-editing context.
GitHub — letta-ai/letta