Most agent evals grade task outputs. HarnessDev grades something we usually treat as fixed scaffolding: the execution harness itself — can the model build and then improve the infrastructure it runs inside?

This matches what I see building agentic systems: model weights are rarely the bottleneck — routing, retries, tool wiring, and verification in the harness are. If an agent can measurably evolve its own harness but only for the model that grew it, we’re drifting toward per-model harnesses rather than one portable framework. The HF paper page has the early reactions. Would you ship a harness your agent tuned itself, knowing it only holds up under the exact model that wrote it?

tags: [ agentic-ai ] [ llm-ops ] [ research ]