“No single memory substrate wins” reads like a hedge until you see what the arXiv evaluation actually measured: the same design choice that helps one workload actively hurts another, inside the same agent.

This maps onto a problem I keep hitting: teams treat agent memory as one component — “add a vector store” — when the HF paper page argues that memory is a routing decision, not a fixed dependency. The retrieval stack that makes your RAG QA sing can quietly wreck a long-horizon planning agent by burying the one piece of state that mattered.

If substrate routing really is necessary, what’s the routing signal — task type, horizon length, or something the agent has to learn online?

tags: [ agentic-ai ] [ rag ] [ llm-ops ] [ research ]