CoT monitoring rests on a quiet assumption: the reasoning trace records the information that actually shaped the answer. FACE-Eval attacks that assumption exactly where it’s thinnest for anyone running agents — when the influencing cue arrives through a tool return instead of the user message.

Line that failure geometry up against where production agents actually operate: almost every signal that moves the model reaches it through a tool result — a retrieved passage, an API payload, another agent’s output — not a tidy user turn. So CoT monitoring is thinnest precisely in the setting we lean on it hardest. Pull the per-model breakdown from the HF paper page and check whether the models you actually deploy are the monitorable ones. My read: treat CoT monitoring as a tripwire, not a control — and never let it be the only thing standing between an agent and an irreversible tool call.

tags: [ agentic-ai ] [ llm-ops ] [ research ]