The number that jumps out of HarnessRisk: attack success ranges from 12.6% to 80.9% across the same models. The variance lives in the harness and its configuration, not the weights. If you run agents in production, that spread should unsettle you more than any single model’s safety card.

The paper’s move is to stop treating agent safety as one thing and split it across six lifecycle phases, then embed an adversarial instruction inside an untrusted workflow artifact for each. That’s the failure mode I actually see: not a model deciding to be evil, but a tool result or a stored scrap of state carrying an instruction nobody sanitized.

Most agent safety dashboards I’ve seen measure “did the model recognize this was bad” — and a model that recognizes danger while still executing the dangerous action is scoring well on the wrong metric. The HF paper page has the sandboxed case breakdown.

If you’re running agents with tool access and persistent state, which of your six phases have you actually red-teamed — or are you assuming the model’s refusal is your control plane?

tags: [ agentic-ai ] [ llm-ops ] [ research ]