Most code-retrieval evals I’ve inherited measure whether the embedding surfaces “code about the same topic.” That’s the wrong thing to measure once the retriever is feeding a coding agent — what matters is whether the top-1 result actually works. ExecRetrieval makes that failure legible.

The trick is the search pool. For each of the 939 Python tasks, the authors plant one canonical implementation and up to four execution-verified buggy variants generated by single-edit mutation of that same file. If the embedder is doing topical clustering rather than functional discrimination, the correct answer ranks below its own near-clones. Under paired McNemar tests across 23 dense embedding configurations plus BM25, the top hosted system hits exec@10 = 1.00 but only exec@1 = 0.331. On the four leading systems, the rank-1 miss is a paired buggy variant 91.5-99.4% of the time, and the canonical scores below at least one of its four distractors in 67-78% of queries.

That is the whole story I care about for production RAG-over-code: a retriever that looks tuned on lexical benchmarks may be a coin flip once you plant execution-negative distractors. The exec@10 number would still make the retriever’s slide look great, and the reranker downstream would then have to do all the actual discrimination work under a latency budget it wasn’t sized for.

The HF paper discussion links the dataset, the execution oracle, and the pairwise test artifacts, which is what you actually need to reproduce this on your own retriever.

If your retriever handed a coding agent the correctly-shaped buggy version 70% of the time, how long would it take you to notice from downstream eval alone?

tags: [ rag ] [ llm-ops ] [ research ]