The quiet failure in production long-context RAG is that one agent reads the document front to back, dragging a compact memory along as it goes. That couples two things that have no business being coupled: how far it has read and how hard it is reasoning. Latency scales with document length, and accuracy swings depending on where the answer happens to sit.

PARSER breaks that coupling. A bank of frozen subagents each own a single chunk and read in parallel; a lead agent runs scatter-gather rounds — broadcast a query, aggregate the evidence, ask a deeper follow-up conditioned on what came back. Only the lead agent is trained with RL. The readers stay off-the-shelf.

What makes this worth a second look isn’t the benchmark delta. It’s that the lead agent’s scatter-gather loop is iterative retrieval wearing an agent costume: a learned query planner over frozen chunk readers, which is most RAG stacks already. If your long-context eval measures multi-hop recall and P99 separately, you can tell whether the win is accuracy, latency, or both before committing to the architecture. The HF paper page has the ablations.

The part I’d want stress-tested before shipping: what happens when the frozen subagents return confident, contradictory evidence? A learned lead can plan queries, but nothing here trains it to arbitrate a disagreement it didn’t know to look for.

tags: [ rag ] [ agentic-ai ] [ research ]