Your first responder for every production incident

Adoe handles the first response to every production alert. It investigates, fixes known problems with your approval, and calls a human when judgment is needed, with the evidence attached.

Every workspace starts read-only. Adoe changes nothing until you turn actions on.

Example incident #4821checkoutprod

Investigating Resolved
  1. Alert received

    checkout_error_rate from Grafana: 7.3% of requests failing

  2. Grouped into one incident

    3 related alerts in 15 minutes became 1 incident

  3. Investigated with read-only access

    Deploy v2.48.1 shipped 8 minutes earlier. Errors only on checkout.

  4. SOP matched, checks passed

    checkout-post-deploy-errors picks runbook rollback.yml

  5. Approval requested

    Posted to #checkout-oncall with the evidence

  6. Approved by on-call

    In Slack, 35 seconds after the request

  7. Runbook ran

    GitHub Actions rollback.yml deployed v2.48.0

  8. Verified, then resolved

    Errors at 0.2% on checkout's own series

Example incident

Works with the tools your on‑call team already uses

  • Grafana
  • PagerDuty
  • Sensu
  • Splunk
  • Slack
  • GitHub
  • AWS
  • Uptime.com
All integrations and their status

What you will know after one week

During evaluation, adoe runs in shadow mode on your own alerts. It investigates and decides, but changes nothing. Every number below comes from your own data.

Your noisiest alerts

Which checks fire most, and which ones your team marked as false alarms. Adoe can propose a tuning change as a pull request.

Source: alert analytics over 30 days, and the false-alarm button on each alert.

The fixes it would have run

Each week adoe posts a report to Slack. It lists the fixes it would have run, and how many of those alerts your team later resolved.

Source: the weekly shadow report. It counts decisions that passed every safety check except your workspace setting.

Whether its confidence holds up

Every day, adoe compares its confidence with what actually happened. You see whether it is earning more autonomy before you give it any.

Source: the Learning health score (calibration), from outcomes your operators confirm.

How it works

Every alert goes through the same three steps. A person stays in control of the step that changes things.

See every stage, with its inputs and safety checks
  1. Step 1

    Investigate, read-only

    Adoe groups related alerts into one incident. It reads recent deploys, the alert's definition, metrics and logs. Then it states a likely cause with a confidence score.

    Show the stages
    • Receive. Alerts arrive from your monitoring tools. Repeats are dropped.
    • Group. Related alerts from the same service within 15 minutes become one incident.
    • Gather evidence. Recent deploys, the alert's definition in your repository, metrics from Grafana, logs from Splunk.
    • Form a hypothesis. A likely cause, a confidence score, and the impact scope: what else the problem could affect.
  2. Step 2

    Act, with your approval

    Your SOP checks the evidence and picks the runbook to run. A person approves in Slack or in the dashboard before anything runs.

    Show the stages
    • Check. An SOP is your procedure for one kind of alert. Its conditions must match before it can pick a runbook.
    • Decide. Adoe chooses to run, recommend or escalate, based on its confidence.
    • Approve. The request shows the evidence, the target and the impact scope.
    • Run. A runbook is the steps that make the fix. Adoe runs yours through GitHub Actions, AWS Systems Manager, a script or SSH.
  3. Step 3

    Verify, or escalate

    Adoe checks the alert's own signal before it resolves anything. If it cannot confirm the fix, it escalates to a person with the evidence.

    Show the stages
    • Verify. Adoe reads the metric for the alert's own host or service.
    • Resolve. Only when the check passes.
    • Escalate. The hypothesis and evidence go to the alert's Slack thread.
    • Learn. Your team's feedback changes how much adoe trusts each SOP.

Safe on production by default

The hard part of automation is knowing when not to act. These four controls apply to every workspace.

Read-only by default

Investigation uses read-only tools. A new workspace cannot change anything until you turn actions on.

Approval for every change

Each action waits for a person to approve it. Adoe's own actions can never delete, terminate or destroy resources.

Verified before resolved

After a fix, adoe checks the alert's own signal. If it cannot confirm recovery, it escalates instead of closing the incident.

Every step on the record

Each investigation, decision, approval and action is saved on the incident, with who approved it and when.

Read the security overview

See adoe work on your own alerts

Book a technical demo. We show adoe investigating an incident against your own SOPs, and where it stops for a person.