Most agent security benchmarks share a quiet weakness: the environments are hand-built and the injection points are decided in advance. ToolHazard argues that’s exactly what makes them flattering — an agent can only get caught where the benchmark author thought to plant the attack.

Its answer is to generate the adversary instead of scripting it. An environment simulator synthesizes executable, stateful tool environments; an attacker agent hunts for viable injection points and writes environment-specific payloads; a user simulator supplies benign long-horizon tasks so the attack has to survive a real workflow. The resulting ToolHazard-Bench stress-tests agents against indirect prompt injections embedded in tool outputs rather than in the user’s prompt — which is where the dangerous ones live once your agent starts acting on state it didn’t write.

The finding that should change how you run evals: injection timing and placement materially move attack success. A static benchmark with fixed injection slots hands you one optimistic number and hides the variance. If your guardrail clears a single hand-authored scenario, you’ve learned it survives that scenario, not that your tool-using agent is safe. The other result worth noting is that alignment data generated by ToolHazard improves security on both its own bench and AgentDojo while preserving benign task utility — the synthetic adversarial environments transfer, which is what makes this worth wiring into a pipeline rather than reading once.

The HF paper page positions this as scalable security research, and the production reading is blunter: your injection test set is a coverage problem, and coverage is something you generate, not something you curate by hand. If attack success depends on where and when the payload lands, what does a single passing red-team run actually prove about your agent?

tags: [ agentic-ai ] [ llm-ops ] [ research ]