DeepSeek shipped V4-Flash today, and the headline isn’t the price — it’s that the cheap “Flash” tier reportedly beats their own V4-Pro-Preview on agent benchmarks. That inverts the usual tiering, where a flash/mini model trades capability for latency. Here the release notes claim the opposite on exactly the tasks that matter for tool use.

The pattern worth watching is harness-native model releases. A model “specifically adapted for Codex” is a bet that the agent scaffold, not the raw weights, is where the next capability jump lives — which lines up with what I see running agentic workflows in production, where the harness often decides more than the checkpoint does. The HN discussion is already picking apart whether “minimal mode, max effort” is a fair way to report a benchmark.

If a flash-tier model can top a pro tier on agent evals, is the pro/flash split about capability anymore — or just about who pays for the extra tokens the harness burns?