Skip to content

Quality (dw-quality)

The Quality agent monitors your data across five dimensions — completeness, accuracy, consistency, freshness, and uniqueness — and rolls them into a weighted 0–100 score per dataset, with the breakdown and trend visible so you know why a score moved. Anomaly detection is statistical (z-score against 14-day baselines), and detected anomalies are deduplicated hard: 50–100 raw signals typically collapse to 5–10 actionable ones, because an alert feed nobody reads is worse than no alert feed.

Its operating discipline matters as much as its math. It re-verifies a failure before alerting on it, distinguishes “the data is wrong” from “the check is wrong”, and never silently corrects, filters, or imputes data to make a check pass — any correction goes through the governed write path with human approval. Genuine breaks are handed to the Incidents agent rather than absorbed into monitoring.

  • Full quality profiling. run_quality_check profiles null rates, uniqueness, distributions, referential integrity, freshness, and volume for a dataset — all columns or a chosen subset.
  • A score you can interrogate. get_quality_score returns the 0–100 score with its per-dimension breakdown and trend, not just a number.
  • Deduplicated anomaly detection. get_anomalies lists detected anomalies classified by severity (critical / warning / info), deduplicated to the actionable set by default.
  • Quality SLAs. set_sla defines metric thresholds with severity levels per dataset; violations are designed to trigger alerts within a minute.
  • Tests born with the pipeline. create_quality_tests_for_pipeline generates quality test specs for a pipeline — this is what the Pipelines agent calls so new pipelines arrive with tests.
  • Estate-level summary. get_quality_summary aggregates quality across datasets so you can see where the estate stands, not just one table.

“Run a quality check on analytics.orders — all columns.”

“Why did the quality score on the customers table drop this week?”

“Show me critical anomalies from the last 48 hours — deduplicated, not the raw feed.”

“Set an SLA on finance.revenue_daily: null rate under 1%, freshness under 6 hours, critical severity.”

  • Warehouses and lakehouses — Snowflake, BigQuery, Databricks for profiling real tables.
  • Quality suites — Great Expectations, Soda, Monte Carlo results flow in through the connector gateway.
  • dbt — test results as a quality signal.

See the connector catalog for setup.

The agent starts in 🟡 Evaluation on built-in sample data — the profiling, scoring, and z-score anomaly algorithms are the real thing, run on a realistic sample estate. It earns 🟢 Connected per system through a passing live test. See Verify your setup.

  • In 🟡 Evaluation, scores and anomalies describe the sample estate, not your warehouse — useful for judging the workflow, meaningless as a statement about your data.
  • Anomaly detection needs a baseline: after connecting a new system, expect the 14-day baseline to build before statistical detection is at full strength.
  • The agent will not fix data for you silently. Corrections are explicit, traced, and human-approved — if you want a number quietly adjusted, this is the wrong tool.
  • A failed check is a signal, not an alert. The agent re-verifies before alerting, which trades a little latency for a lot less noise.