Systima put Claude Code and OpenCode on the same model, the same machine, and the same tasks, then spliced a logging proxy at the API boundary to measure exactly what each harness spends before your prompt even arrives. The headline gap is real, but the nuance is where the operating lesson lives.

The takeaway isn’t “harness X is bloated.” It’s that every scaffolding token is working context you can’t spend on the actual code, and cache stability matters more than prompt size. If you run agents under audit — EU AI Act Article 12 expects you to log what your system actually does — “what does my agent send” should be answerable from data, not folklore. The full measurement writeup shows the method; the HN discussion argues over whether cache reads make the baseline moot. When did you last read your agent’s payload at the API boundary?

tags: [ agentic-ai ] [ llm-ops ] [ ai-infrastructure ]