The headline number in Qwen-Planner-Agent — best overall on MobilePA-Bench — is the least interesting thing in it. What earns a read is the “action-feedback-verification contract” that treats the planner model and its harness as one system that improves together, instead of two teams lobbing releases over a wall.

The HF paper page frames this as model–harness co-evolution, and that’s the claim I’d pressure-test first: does the contract hold on a messier tool surface than a phone screen, or does co-evolution just relearn the harness every time the base model moves? If you’re running agents in production today, would you spend the next quarter on a better base model or on making your harness improvable at all?

tags: [ agentic-ai ] [ research ] [ llm-ops ]