Two sibling models that score 0.660 and 0.650 on the same tool-use metric look like they have the same capability. Keyword Harnesses Fail Open shows one of them emits valid tool calls on 6 of 6 training prompts and the other on 0 of 6 — the benchmark simply can’t tell them apart.

This is the quiet failure in every in-house agent eval I’ve shipped: the harness measures whether the output looks like a tool call, not whether the model can produce one on a prompt it hasn’t memorized. The HF paper page lays out the full four-level ladder. If you’re scoring agent tool use with substring or keyword matching today, when did you last diff your eval’s “passes” against what the model literally emitted?

tags: [ llm-ops ] [ agentic-ai ] [ research ]