Most generative-retrieval papers quietly assume you’ll tear out your retrieval stack and rebuild around a new index. CoGR is interesting because it refuses to. It trains LLMs to emit compact keyword sets on both the query and item side, then matches them through a plain inverted index — the sparse infrastructure you already run.

The compatibility story is the real pitch. If you run hybrid search in production, a technique that upgrades the sparse leg without touching your index plumbing is a far easier sell than another embedding rebuild. The HF paper page has the training curves. So here’s my open question: how much of that 36% survives when the item universe churns daily instead of sitting still?

tags: [ rag ] [ ai-infrastructure ] [ research ]