Most agent benchmarks grade the finished agent. τ^τ-bench grades the agent that has to build the agent — and the gap it exposes is the one I actually care about in production.

What makes this land is that none of those failures are model-capability gaps. They’re engineering-judgment gaps: interrogate the data, clarify the ask, tune the cost budget. That’s the daylight between a demo agent and a deployed one, and it’s exactly what the HF paper page frames as the measurable target. If the frontier stalls at ~24% here while raw coding benchmarks keep climbing, does that tell us the models are the bottleneck — or that we’ve been benchmarking the wrong half of the job?

tags: [ agentic-ai ] [ llm-ops ] [ research ]