We pour a lot of effort into hand-tuning agent prompts, tools, and workflows, then freeze the scaffold around the model and never touch it again. Hierarchical Self-Improvement asks what happens if the harness itself becomes the thing that learns — while the model stays frozen.

What I like is the framing: the harness is a tunable surface, not a fixed artifact you ship once. For production agents that’s the cheaper axis — you can evolve orchestration, retries, and tool wiring against real environment feedback without retraining anything. The HF paper page links the code. The catch is feedback fidelity: most production tasks don’t hand you a clean reward signal, so the real work isn’t the evolver — it’s building an environment honest enough to evolve against.

tags: [ agentic-ai ] [ llm-ops ] [ research ]