From alert to verified fix, in six stages

Every alert goes through the same stages. Each one has a defined input, a defined action, a safety check and a result. Adoe stops for a person before anything changes.

01

Receive

Alerts arrive and become one incident.

Input
Alerts from Grafana, PagerDuty, Sensu, Splunk, Uptime.com and Railway. Each tool sends them to a private webhook address for your workspace. Adoe also checks Grafana, PagerDuty and Uptime.com on a schedule, in case a webhook is missed.
Action
Adoe turns every alert into one common format. It drops repeats of the same alert within 30 minutes. It groups related alerts from the same service and environment within 15 minutes into one incident.
Safety
Read-only. The webhook address contains a secret token. PagerDuty signatures are checked when you set a signing secret. Requests over 2 MB are rejected, and each source has a rate limit.
Result
One incident, not a page for every alert.
02

Investigate

Adoe gathers evidence and forms a likely cause.

Input
The incident, plus context: recent deploys, the alert's definition in your repository, the live check definition from your monitoring tool, metrics from Grafana and logs from Splunk.
Action
Adoe reads the evidence and forms a hypothesis: a likely cause, a confidence score, and the impact scope. Impact scope (sometimes called blast radius) is what else the problem could affect. When no SOP matches, a diagnostic agent looks further with read-only tools and proposes one action.
Safety
Read-only tools only. Each investigation has a two-minute limit. The diagnostic agent is limited to 8 tool calls and 30 seconds. If a source does not answer in time, adoe skips it and carries on.
Result
A hypothesis with confidence and impact scope.
03

Decide

Your SOP checks the evidence and picks a runbook.

Input
The hypothesis and your SOPs. An SOP is your procedure for one kind of alert: what to check, which runbook to run, and how to confirm the fix.
Action
Adoe finds the SOPs for this alert and checks their conditions, such as environment or service. Then it chooses one of three actions: run the runbook, recommend it to a person, or escalate.
Safety
  • Below a confidence of 0.8, adoe escalates instead of recommending. Each SOP also sets its own, usually higher, bar for running without a person.
  • Adoe never runs a fix it cannot verify. If the SOP's check needs a metric source that is not connected, adoe recommends instead.
  • Your workspace setting decides how far adoe may go: off, shadow, approve or auto. Every workspace starts at off.
Result
A decision with its reasoning, saved on the incident.
04

Approve

A person decides before anything changes.

Input
The proposed runbook, the target, the impact scope and the evidence.
Action
Adoe posts an approval request in the alert's Slack thread. You can also run it from the dashboard, where you see a preview first and can do a dry run.
Safety
  • In approve mode, nothing runs until a person approves it.
  • Adoe's own actions can never delete, terminate or destroy resources. Actions that can add cost, like starting or scaling, always need approval.
  • Approval requests for adoe's own actions expire after 15 minutes. A recommendation that nobody acts on escalates after 24 hours.
Result
An approved action, with who approved it and when.
05

Act and verify

The runbook runs. Adoe checks that it worked.

Input
The approved runbook. A runbook is the steps that make the fix: a GitHub Actions workflow, an AWS Systems Manager document, a script, or an SSH command.
Action
Adoe runs the runbook's steps in order, with the alert's details filled in. Then it checks the result the way your SOP says: a metric, an HTTP endpoint, or a command.
Safety
A metric check reads only the series for the alert's own host or service. If adoe cannot get data, cannot tell which series is the right one, or the check does not pass in time, it escalates. It never resolves an incident it could not verify.
Result
A verified fix, or an escalation.
06

Escalate and learn

A person gets the full picture. Next time is faster.

Input
Anything adoe could not verify, could not match to an SOP, or was not confident about.
Action
Adoe posts its hypothesis and evidence in the alert's Slack thread and follows up there. For an alert with no SOP, it drafts one for your team to review. Your team marks each outcome as fixed by the SOP, fixed by hand, or a false alarm.
Safety
A person decides what happens next. SOPs that adoe drafts, or imports from GitHub or Confluence, only run after an admin approves them.
Result
A person who knows what is going on within seconds. And an SOP for next time.

Terms used on this page

Alert
What your monitoring tool sends when a check fails.
Incident
What adoe works on: one or more related alerts, grouped together.
SOP
Your procedure for one kind of alert. It checks the conditions, picks the runbook, and says how to confirm the fix. An SOP decides.
Runbook
The steps that make the fix, such as a GitHub Actions workflow or a script. A runbook acts.
Impact scope
Everything a problem, or a fix, could affect. Also called blast radius.
Shadow mode
Adoe investigates and decides, but changes nothing. Each week it reports what it would have done.
Escalate
Hand the incident to a person, with the hypothesis and evidence attached.

See each stage on a real incident

Book a technical demo. We walk through an incident end to end, including where adoe stops for a person.