The interesting claim in NeoHorse-1 isn’t “recursive self-improvement” — that phrase is doing a lot of marketing work across the field right now. It’s the mechanism underneath: the routing layer, the thing that decides which model in a heterogeneous pool handles each turn, becomes the training signal itself.

What lands for me is the reframe of routing from a cost knob into a data-generation engine. Most of us treat the router as the thing that saves money at inference time; NeoHorse treats every routing decision as a labeled training signal. Whether that survives past one iteration — before a model starts routing in ways that reinforce its own blind spots — is the question the HF paper page leaves open. My bet: the failure mode of self-routing curricula won’t be raw capability, it’ll be diversity collapse.

tags: [ agentic-ai ] [ llm-ops ] [ research ]