Sylvain Kalache’s piece names something most AIOps dashboards will happily hide: your average MTTR can fall while your worst incidents get slower. AI clears the routine cases — inspect alerts, form hypotheses, correlate deploys, ship the fix — and the routine cases are exactly how engineers built the intuition they need for the incident automation can’t touch.

The production question isn’t “should we let AI run incidents” — for routine ops that’s already the default. It’s whether your observability and on-call design deliberately keep humans engaged enough to stay sharp for the tail. The HN discussion is full of engineers who’ve already felt that gap widen. If your agent resolves 95% of incidents, who on your team is still practicing for the 5% that pages the CEO?

tags: [ llm-ops ] [ agentic-ai ] [ industry ]