The useful takeaway from this empirical study of coding-agent harnesses is that there’s no single good harness. Planning, action space, and context management each pay off differently depending on which model you’re driving and the token budget you’re driving it under. That’s an inconvenient result if you’ve been treating your harness as a platform you build once and reuse across every model.

A few findings that map straight onto how I think about running agents in production:

That fourth point is the uncomfortable one for anyone who spent months curating a tool catalog: if your model is strong enough, a chunk of that orchestration is cost you’re paying for no accuracy. The HN discussion is already arguing whether the context-management gains are a design effect or just more token spend — the right thing to check before you copy any of these defaults into your own framework.

So how much of your agent scaffolding is genuinely buying accuracy, and how much is expensive habit your current model has already outgrown?

tags: [ agentic-ai ] [ llm-ops ] [ research ]