Most of us fix flaky multi-step tool calls with more scaffolding — retries, a planner, a critic, a bigger model. This paper points at a different lever: the model’s depth of computation at inference, through looped (recurrent) layers, is what makes compositional tool chains hold together.

The distinction that matters in production is where the gains land. Recurrent computation helps most on compositional, dependency-aware tool use — coordinating multiple API calls, carrying intermediate state, preserving dependencies across hops — and barely moves isolated single-call invocation. That maps onto exactly where real agents fall over. A one-shot function call is easy; it’s the third tool call that depends on the parsed output of the first that breaks.

If variable recurrent depth is a genuine dial for tool-call reliability, the uncomfortable question is how much of the planning logic we’ve pushed into orchestration frameworks belongs back inside the model. The HF paper page is worth watching for whether that adaptive-depth trade-off survives outside these three benchmarks.

tags: [ agentic-ai ] [ ai-infrastructure ] [ research ]