Anyone who has shipped an agent knows the uncomfortable part: most of the reliability lives in the harness — the prompts, tool configs, and control logic wrapped around the model — not in the weights. AutoSaddler treats that harness as code and optimizes it offline from failure traces, and the framing is more useful than the benchmark deltas.

The part I keep circling back to: harness tuning today is manual, expensive, and mostly tribal knowledge — the thing a senior engineer does by staring at traces late at night. If “harness improvement as an offline learning problem” holds up outside these benchmarks, the shift that matters isn’t better scores; it’s that harness engineering stops being a craft and starts being a pipeline. The HF paper page has the ablations, and my bet is that diagnosis quality, not the patch generator, is what actually carries the gains.

tags: [ agentic-ai ] [ llm-ops ] [ research ]