The finding here should make anyone shipping a voice agent pause: the retrieval tricks we add to improve multi-hop answers can make the system less robust once a speech recognizer sits in front of it.
- ๐ฏ Structure amplifies upstream error. Entity-graph linking and iterative reformulation raise absolute F1, but the arXiv paper shows the clean-vs-noisy F1 gap widening 36โ67% versus naive dense retrieval across HotpotQA, 2WikiMultiHopQA, and MuSiQue.
- ๐ Entities are the fault line. Corrupted query entities drive 87โ96% of the degradation on 2WikiMultiHopQA โ exactly the tokens ASR mangles most, and exactly what graph hops then propagate.
- โ ๏ธ Absolute score hides the risk. A config that wins on clean text can be the one that falls hardest under accented speech; the leaderboard number wonโt tell you which.
- ๐ Test the whole pipe, not the retriever. They synthesize four accents through neural TTS to vary word-error rate โ a pattern worth stealing for any conversational-AI eval harness.
- ๐ก Surface-form patches arenโt enough. Two lightweight mitigations left most of the gap intact, which points the fix upstream โ entity-aware ASR, confidence-gated hops โ rather than at the retriever.
If your RAG eval only ever sees clean, typed queries, youโre grading the easy half of the problem. Worth reading against the HF paper page before you assume that more retrieval structure is strictly safer.
tags: [ rag ] [ conversational-ai ] [ knowledge-graphs ] [ research ]