When the executive who’s been most aggressive about agents tells Reuters that agent development is going slower than expected, that’s worth more than another demo. The gap between a single agent nailing a scripted task and a fleet of them running unattended in production is where most timelines quietly die. Here’s what actually eats the schedule, from the side that ships this:

None of this means agents don’t work. It means the demo-to-production tax is real and mostly invisible until you’re paying it. The HN discussion has the usual split between “told you so” and “still early,” and both can be right.

My bet: the teams who slip quietly on eval infrastructure now are the ones who ship reliable agents first. Who’s actually measuring the ten-step failure rate?