Anyone who has run a prompt optimizer over a multi-agent pipeline has felt this failure: a tweak that clearly improves an agent’s answers also breaks the run, because the same prompt that generates content also carries the routing rules, output schema, and termination signal the surrounding code depends on. This paper names the problem cleanly — the prompt is doing two jobs, and the optimizer can’t tell them apart.

The HF paper page frames this as prompt optimization, but the deeper move is treating agent orchestration like a type system: contracts the optimizer isn’t allowed to touch, content it is. If your multi-agent framework still smuggles control flow through free-text prompts, how many of your “the model regressed” incidents were actually the optimizer editing your routing by accident?

tags: [ agentic-ai ] [ llm-ops ] [ research ]