We keep calling the human approval prompt the last line of defense against a rogue agent. Scale X’s data from 40,000 runs of a browser game — where you play the human approving an AI coding agent’s commands under time pressure — suggests it’s a thin line. Mean accuracy was 66.3%. The reviewer missed roughly one threat in three.

The miss-rate breakdown is the useful part, because it tracks how obvious the damage is, not how dangerous it is:

The quiet exfil and credential reads — the ones that actually end incidents — are exactly where human attention falls off. And this was a game where 34% of commands were threats and players knew they were being tested. In production the base rate is a fraction of that, which makes vigilance worse, not better: approval fatigue sets in fast when 99 of 100 commands are benign.

The operational takeaway for anyone shipping agents: human-in-the-loop is a backstop, not a control. Allowlists, static command classification, and scoped credentials have to catch the exfiltration and scope-violation classes before the human ever sees them. The HN discussion has the usual split between “just read every command” and people who’ve felt the fatigue firsthand.

If a click-through approval catches two threats in three, is it a guardrail or a liability-shifting UI?

tags: [ llm-ops ] [ agentic-ai ] [ enterprise-ai ]