The detail worth sitting with in Simon Willison’s writeup isn’t that a new model is capable. It’s that he asked it to inspect a dependency for a CSS scrollbar bug, walked away, and came back to find it driving a real browser he never told it to open. Proactiveness is a capability and a control surface at the same time — and most agent harnesses only budget for the first.

The HN discussion splits between people who love this and people unnerved by it, which is the right tension.

My prediction: the next year of agent engineering is less about raising capability and more about making restraint a first-class, testable behavior. Which would you rather ship to production — the agent that does too much, or the one that asks first?