Everyone shipping GUI agents already knows they’re brittle. The useful result in AnTrap isn’t that 16 leading models all crumble under runtime anomalies — it’s that the failures split cleanly into ones you can train away and ones you can’t.

The benchmark injects perturbations into live agent trajectories — pop-ups, action misuse, dead states — organized into a four-layer taxonomy (State, Thinking, Action, Round) with ten subcategories, while keeping every task still solvable. Then it runs GRPO training in both clean and adversarial environments to see which failures the model can actually learn out of. That last step is what makes this more than another “agents are fragile” leaderboard.

The framing maps directly onto the retries-and-fallbacks work anyone running production agents already does: some anomalies deserve adversarial training data, and some deserve a guardrail that detects the deadlock and hands control back. The HF paper page has the full taxonomy. So which of your agent’s failures are you still trying to fine-tune away when they’re really a recovery-design gap?

tags: [ agentic-ai ] [ llm-ops ] [ research ]