Most eval harnesses report one number per benchmark, and that number quietly conflates two different things: whether the model is right, and whether it’s stable. The reframing in this paper — generalization is stability, not accuracy — changes what you instrument, not just what you report.

The formal objective is in the arXiv paper; the HF paper page is the faster way in. If you re-ran your last eval with five paraphrases of every prompt and scored the variance, how many of your “passing” models would still clear the bar?

tags: [ research ] [ llm-ops ] [ enterprise-ai ]