The headline from Fernando Irarrázaval’s hackmyclaw experiment is that 6,000+ emails from 2,000 people failed to pull a single secret out of his OpenClaw agent. The more useful finding for anyone running agents in production is the inverse: the model held, and almost everything around it broke.

This maps straight onto production agent work: the confused-deputy risk is real, but the tickets you’ll actually file are context bleed, provider fraud-flags, and memory hygiene — not a clever subject line. The HN thread splits between “this proves injection is overblown” and “one-shot email is the easy case.”

My bet is the second camp: hand the same agent 20-email back-and-forth threads instead of one-shots and the success rate stops being zero. One-shot injection is the spelling test; multi-turn is the exam. Where has your agent’s boundary actually been probed — single messages, or sustained conversations?