LLM-as-judge is load-bearing in most eval pipelines now, and this paper names a failure mode a lot of us have felt but never measured: the judge rewards how an answer is written, not what it actually says.

Swap “manuscript” for “support ticket” or “agent trajectory” and this is every production LLM-judge I’ve shipped — the judge quietly tracks fluency, and your offline metric drifts from what users actually care about. The HF paper page has the breakdown. If your eval judge scores a verbose answer above a terse correct one, you’re measuring rhetoric too — do you know by how much?

tags: [ llm-ops ] [ research ]