Real-SWE (benchmark page) does the one thing public coding benchmarks structurally can’t: it runs agents against private production codebases whose code and fixes were never on the internet to train on. The leaderboard is sobering, but the failure mode is the real story.

The HN discussion is worth a read, though the obvious critique — can 18 codebases generalize? — matters less than the failure mode, which won’t move much with sample size. My bet: the next leap in enterprise coding agents comes from better context assembly over private repos, not bigger models. Which will your team invest in first?

tags: [ agentic-ai ] [ enterprise-ai ] [ llm-ops ]