Every coding-agent leaderboard I’ve looked at was built on SWE-bench-style prompts: long, formal, information-rich problem statements distilled from curated GitHub issues. RealSWE points out how little that resembles what users actually type — and then shows the gap is wide enough to reorder the ranking.

The measurement I care about is the distribution mismatch. Problem-statement-only requests are 88% of real prompts but 7% of benchmark problems; 87% of real prompts are casually written while 94% of benchmark problems are formal. When the authors rebuild 381 task families from SWE-bench Verified and Pro to match the real distribution — same underlying task and gold patch, only the information composition and style vary — resolution rates fall 6.4 points on average and, more tellingly, model rankings change. If you picked your agent because it topped a formal benchmark, that ranking may be an artifact of prompt shape.

The actionable part is sharper than the headline. Two information categories carry the signal: stating the Desired Behavior and the Motivation. Environment Information and Reproduction Steps — the fields we habitually pad issue templates with — mostly add tokens without measurable benefit, and linguistic style barely registers. So the lever isn’t a fancier model; it’s getting the intent stated, and 88% of real prompts arrive without it.

The HF paper page has the taxonomy and the per-category ablations, which is what you’d need to run this shape of eval against your own harness.

If your agent’s benchmark rank was set on formal prompts your users never write, how confident are you it holds on the casual one-liners they actually send?

tags: [ agentic-ai ] [ llm-ops ] [ research ]