When the AI SRE Fumbles: The Deskilling Trap Hitting Incident Response
As autonomous 'AI SRE' agents absorb routine incidents, engineers lose the practice that sharpens them for rare SEV0s. Sylvain Kalache's widely shared essay revives Bainbridge's 1983 'Ironies of Automation' and calls for aviation-style incident simulators on every on-call rotation.
The pitch for AI SRE tools is seductive and, by most accounts, increasingly real. Modern incident-response agents inspect the alert, form hypotheses, query telemetry, correlate the blast radius with recent deployments, and — in the bolder deployments — implement the fix themselves. “These tools do it all,” writes Sylvain Kalache, Head of AI Labs at Rootly, in an essay published September 4 that hit the front page of Hacker News this week. “As much as I love to see it, I have a major concern: we are losing touch with our systems.”
That sentence is the whole argument in miniature. Kalache is not an AI skeptic — he spent years building these capabilities, and his essay opens with a confession: as an SRE at LinkedIn in 2012, he designed a self-healing system that learned from previous incidents. AI capabilities were nowhere near what we have today, and it stayed a prototype — but the vision is now shipping as products. His worry is not that the tools fail at routine work. His worry is that they succeed at it.
The paradox is 43 years old
The counterintuitive core of the essay is borrowed, openly and deliberately, from human-factors researcher Lisanne Bainbridge. Her 1983 paper “The Ironies of Automation” remains the definitive statement of a paradox every operator eventually meets: automation removes exactly the routine work that lets humans safely build intuition for how a system behaves and fails, while leaving humans responsible for precisely the novel, abnormal situations automation cannot handle.
Bainbridge’s conclusion was bracing for its time and reads as prophetic now: operators of automated systems need to be more skilled and more trained than before, not less — because their remaining work is the hard part. Software engineering spent four decades mostly dodging this iron law because ops was never automated enough to atrophy the muscle. The AI SRE era ends that grace period. When an autonomous agent resolves the nighttime capacity alerts, the pager never wakes you — and you never build the pattern library that only comes from paging through failures yourself.
Kalache’s prediction is specific enough to be falsifiable: mean time to resolution (MTTR) for routine incidents will fall thanks to AI-assisted response, while resolution time for complex incidents will shoot up, because the humans who finally take over are out of practice and struggling to investigate a system they no longer understand from the inside.
The aviation model: rehearse the rare
If automation creates the problem, an older automation-heavy industry points at the solution. Commercial aviation automated most of flying decades ago, and it answered the resulting skill-atrophy risk with mandatory simulation. Under US FAA rules, airline captains return to a simulator every six months for recurrent training or a proficiency check — including scenarios like an engine failure during takeoff. Modern turbine engines fail so rarely (fewer than one in-flight shutdown per 100,000 engine flight hours) that a pilot can complete an entire career without seeing one outside a simulator. The simulator is where the rare emergency is made common.
The essay’s most sobering paragraph is an incident report. On a twin-engine turboprop departure, the right engine’s propeller autofeathered shortly after takeoff. The aircraft was designed to fly on the left engine — the failure was survivable by the book. But the crew misidentified the problem, the aircraft stalled, and it crashed 117 seconds after the first warning. Rare event, degraded hands-on familiarity, catastrophic outcome. “While most software incidents do not threaten lives,” Kalache writes, “that is no reason not to perfect our craft.”
Incident simulators, not incident spectators
The constructive half of the essay argues the software industry needs its own simulators, and that the same LLM technology causing the deskilling can power the training. At Rootly, Kalache’s team has partnered on realistic incident simulations: the engineer takes the incident-commander seat during a simulated e-commerce outage, works real observability tools, and coordinates with LLM-powered stakeholders in Slack — including a demanding CEO and customer support breathing down their neck.
“The result feels real,” he writes. The skills being rehearsed are exactly the ones that decay first: making sense of incomplete information, communicating clearly, coordinating people, and actually running the response rather than watching one.
He explicitly rejects the easy version of the fix. Yes, responders can ask an AI agent to explain the steps it took, the signals it examined, and the evidence behind its diagnosis — and that explanation is valuable. “But explanation and observation are not substitutes for practice,” the essay argues. “You might pick up a few things from watching Serena Williams play, but you only learn tennis by getting on the court, and incident response is no different.” The point lands because Kalache spent over five years building Holberton School around progressive, project-based education — when Dropbox told him its hires were still weak at troubleshooting, his answer was to hand students deliberately broken infrastructure and make them repair it.
Comprehension debt
The essay’s most durable coinage is comprehension debt: a growing gap between how systems actually work and how well the humans on call understand them. It is the technical cousin of what Bainbridge measured in cockpit crews — and it compounds silently, incident by incident that AI resolves without a human in the loop. Tabletop exercises and chaos engineering are not new, the essay concedes; the LLM era just made them mandatory rather than optional. Regular hands-on interaction with the systems you cover, deliberate exposure to unfamiliar failures, pressure practice, and SEV0 coordination drills belong on the on-call readiness checklist now.
The closing irony is Bainbridge’s own, restated for the AI generation: the more successful automation becomes, the less prepared humans may be for the moment it fails.
For platform and SRE leaders, the essay lands at an awkward moment. The same week it spread across Hacker News, vendors were busy marketing fully autonomous operations as the destination. Kalache’s argument does not dispute the destination — he is building toward it himself. It disputes the assumption that humans passively inherit readiness. The aviation industry automated flying and then spent billions on simulators to keep pilots sharp for the 117 seconds that matter. Software, now automating its own operations, has so far spent almost nothing on the equivalent. The “AI SRE” that never sleeps is arriving either way; the open question is whether the human in the loop will still be awake to the system when the AI finally hands the pager back.
Sources
- [1] https://www.sylvainkalache.com/blog/ai-handles-incidents-engineers-lose-touch-with-their-systems
- [2] https://leaddev.com/software-quality/ai-assisted-coding-incident-magnet
- [3] https://rootly.com/blog/how-engineering-leaders-rethink-code-review-in-the-ai-era
- [4] https://dl.acm.org/doi/10.1111/j.1467-8412.1983.tb00341.x