The headline on architect-loop sells an 80% token cut. The part worth copying is the org chart: one model plans and reviews and never writes a line of code, another writes everything and never decides what to build. Single-agent loops drift because the context that wrote the bug is the one grading it — splitting judgment from execution is the fix I keep returning to in production agent work.

The HN thread argues over whether 80% survives contact with a real repo, which is exactly the right thing to argue about. My open question is upstream of the number: when the spec itself is wrong, a reviewer that can’t write code can’t quietly patch a bad gate — it has to bounce the slice back. That discipline is a feature right up until your gate is the bug.