From alert to verified fix, in six stages
Every alert goes through the same stages. Each one has a defined input, a defined action, a safety check and a result. Adoe stops for a person before anything changes.
01
Receive
Alerts arrive and become one incident.
- Input
- Alerts from Grafana, PagerDuty, Sensu, Splunk, Uptime.com and Railway. Each tool sends them to a private webhook address for your workspace. Adoe also checks Grafana, PagerDuty and Uptime.com on a schedule, in case a webhook is missed.
- Action
- Adoe turns every alert into one common format. It drops repeats of the same alert within 30 minutes. It groups related alerts from the same service and environment within 15 minutes into one incident.
- Safety
- Read-only. The webhook address contains a secret token. PagerDuty signatures are checked when you set a signing secret. Requests over 2 MB are rejected, and each source has a rate limit.
- Result
- One incident, not a page for every alert.
02
Investigate
Adoe gathers evidence and forms a likely cause.
- Input
- The incident, plus context: recent deploys, the alert's definition in your repository, the live check definition from your monitoring tool, metrics from Grafana and logs from Splunk.
- Action
- Adoe reads the evidence and forms a hypothesis: a likely cause, a confidence score, and the impact scope. Impact scope (sometimes called blast radius) is what else the problem could affect. When no SOP matches, a diagnostic agent looks further with read-only tools and proposes one action.
- Safety
- Read-only tools only. Each investigation has a two-minute limit. The diagnostic agent is limited to 8 tool calls and 30 seconds. If a source does not answer in time, adoe skips it and carries on.
- Result
- A hypothesis with confidence and impact scope.
03
Decide
Your SOP checks the evidence and picks a runbook.
- Input
- The hypothesis and your SOPs. An SOP is your procedure for one kind of alert: what to check, which runbook to run, and how to confirm the fix.
- Action
- Adoe finds the SOPs for this alert and checks their conditions, such as environment or service. Then it chooses one of three actions: run the runbook, recommend it to a person, or escalate.
- Safety
-
- Below a confidence of 0.8, adoe escalates instead of recommending. Each SOP also sets its own, usually higher, bar for running without a person.
- Adoe never runs a fix it cannot verify. If the SOP's check needs a metric source that is not connected, adoe recommends instead.
- Your workspace setting decides how far adoe may go: off, shadow, approve or auto. Every workspace starts at off.
- Result
- A decision with its reasoning, saved on the incident.
04
Approve
A person decides before anything changes.
- Input
- The proposed runbook, the target, the impact scope and the evidence.
- Action
- Adoe posts an approval request in the alert's Slack thread. You can also run it from the dashboard, where you see a preview first and can do a dry run.
- Safety
-
- In approve mode, nothing runs until a person approves it.
- Adoe's own actions can never delete, terminate or destroy resources. Actions that can add cost, like starting or scaling, always need approval.
- Approval requests for adoe's own actions expire after 15 minutes. A recommendation that nobody acts on escalates after 24 hours.
- Result
- An approved action, with who approved it and when.
05
Act and verify
The runbook runs. Adoe checks that it worked.
- Input
- The approved runbook. A runbook is the steps that make the fix: a GitHub Actions workflow, an AWS Systems Manager document, a script, or an SSH command.
- Action
- Adoe runs the runbook's steps in order, with the alert's details filled in. Then it checks the result the way your SOP says: a metric, an HTTP endpoint, or a command.
- Safety
- A metric check reads only the series for the alert's own host or service. If adoe cannot get data, cannot tell which series is the right one, or the check does not pass in time, it escalates. It never resolves an incident it could not verify.
- Result
- A verified fix, or an escalation.
06
Escalate and learn
A person gets the full picture. Next time is faster.
- Input
- Anything adoe could not verify, could not match to an SOP, or was not confident about.
- Action
- Adoe posts its hypothesis and evidence in the alert's Slack thread and follows up there. For an alert with no SOP, it drafts one for your team to review. Your team marks each outcome as fixed by the SOP, fixed by hand, or a false alarm.
- Safety
- A person decides what happens next. SOPs that adoe drafts, or imports from GitHub or Confluence, only run after an admin approves them.
- Result
- A person who knows what is going on within seconds. And an SOP for next time.
Terms used on this page
- Alert
- What your monitoring tool sends when a check fails.
- Incident
- What adoe works on: one or more related alerts, grouped together.
- SOP
- Your procedure for one kind of alert. It checks the conditions, picks the runbook, and says how to confirm the fix. An SOP decides.
- Runbook
- The steps that make the fix, such as a GitHub Actions workflow or a script. A runbook acts.
- Impact scope
- Everything a problem, or a fix, could affect. Also called blast radius.
- Shadow mode
- Adoe investigates and decides, but changes nothing. Each week it reports what it would have done.
- Escalate
- Hand the incident to a person, with the hypothesis and evidence attached.
See each stage on a real incident
Book a technical demo. We walk through an incident end to end, including where adoe stops for a person.