01Challenge
The platform served users in one hemisphere while the engineering team lived in another. Overnight alerts waited for a person to wake up, and the first hour of every incident went into reconstructing context: which deploy, which service, which similar incident last quarter.
02Solution
An agent that mirrors what an on-call engineer actually does — reads the alert and stack trace, pulls the related deploy diff and prior incidents, proposes a root cause, and opens a draft pull request with a test. Merge stays a human decision.
03Business impact
The overnight rota was retired, incident cost moved from unpredictable engineering hours to a fixed running cost, and senior engineers stopped starting their day inside someone else's outage.