A robot is sprinting at you — do you want it running Claude or Grok? OpenRouter’s Jacky Liang turned that framing into an actual experiment, and the result is a clean argument against trusting leaderboards when you pick an agent model.

The full data is worth a read before your next model-routing decision, and the HN discussion argues fairly about whether a game like this generalizes at all. But the core tension is real: if your eval can’t separate a model that wins from a model that cooperates, which one is your router quietly shipping to production?