OR-Clarify benchmarks the one thing most agent evals quietly assume away: whether the agent knows its instructions are incomplete before it acts. The setup withholds formulation-critical slots from an optimization request and scores the agent on recovering them through bounded interaction with a simulated user — not on solving the problem it was handed. That’s the right frame. In production the request is almost never complete.

The benchmark details live on Hugging Face if you want to run it. What I keep coming back to: we pour enormous effort into making agents answer, and almost none into teaching them to notice when they shouldn’t answer yet.

tags: [ agentic-ai ] [ conversational-ai ] [ research ]