Incidents (dw-incidents)
The Incidents agent is the on-call responder for your data platform. Feed it anomaly signals — a metric that moved, a pipeline that failed, an SLA that slipped — and it classifies the incident into one of six types (schema change, source delay, resource exhaustion, code regression, infrastructure, quality degradation), assigns a severity, and suggests remediation. It works the way a good SRE does: timeline first, then competing hypotheses, then evidence, then escalation.
For root cause, it doesn’t guess from the symptom. It traverses the lineage graph upstream, queries execution logs, and cross-references incident history to build a causal chain with confidence scores — distinguishing the symptom (a missing row count) from the cause (a late-arriving upstream batch, a schema change that broke a join).
Key capabilities
Section titled “Key capabilities”- Incident classification.
diagnose_incidenttakes raw anomaly signals and returns an incident type, severity, and suggested remediation actions. - Graph-based root cause analysis.
get_root_causetraverses lineage up to five or more hops upstream and returns a causal chain with confidence scores, not a single guess. - Playbook remediation with a confidence gate.
remediateruns known playbooks — restart task, scale compute, apply schema migration, switch backup source, backfill data. Automatic execution requires diagnosis confidence above 0.95; anything novel generates a diagnosis report and routes to a human for approval. - Dry-run mode. Simulate a remediation before executing it to see exactly what would happen.
- Incident memory.
get_incident_historyuses vector similarity to find past incidents like the current one, so the third occurrence of a pattern is diagnosed faster than the first. - Structured incident communication. Output leads with severity, phase, and impact surface — which assets, which downstream consumers, since when — in scannable form.
Example prompts
Section titled “Example prompts”“Row counts on
analytics.daily_ordersdropped 40% overnight — diagnose it.”
“What’s the root cause of this incident? Walk the lineage upstream and show me the causal chain.”
“Have we seen an incident like this before? Show me similar past incidents.”
“Dry-run the backfill playbook for this incident before we execute anything.”
Connect it to your stack
Section titled “Connect it to your stack”- Warehouses and lakehouses — Snowflake, BigQuery, Databricks for the signals and the affected assets.
- Orchestration — Airflow and friends, for pipeline run context and restart playbooks.
- Alerting — PagerDuty, Slack, Microsoft Teams, OpsGenie for paging and resolution.
See the connector catalog for setup.
Works before you connect anything
Section titled “Works before you connect anything”The agent starts in 🟡 Evaluation on built-in sample data: a realistic set of incidents, metrics, and lineage you can diagnose end to end before any credential exists. It earns 🟢 Connected per system through a passing live test. See Verify your setup.
Limits, honestly
Section titled “Limits, honestly”- In 🟡 Evaluation, diagnosis and remediation run against the sample estate — no real system is touched until connections are verified.
- Auto-remediation is deliberately narrow: only known patterns above the 0.95 confidence threshold execute without a human. Everything else stops at a diagnosis report and an approval request. That’s a floor, not a temporary restriction.
- Remediation playbooks are a fixed set; incidents outside them route to a human with the evidence gathered so far rather than an improvised fix.
- Root cause quality depends on lineage coverage. If an upstream system isn’t connected, the causal chain stops at the boundary and says so.