Skip to content

Catalog & Context (dw-context-catalog)

The Catalog & Context agent is the swarm’s librarian: it answers “what data do we have, where did it come from, and can I trust it?” It searches your catalog in natural language using hybrid ranking — keyword, graph, and (when enabled) vector signals fused together — traverses lineage upstream and downstream to column level, and assembles complete context for any asset in a single call: schema, lineage, quality, freshness, documentation, and related metrics.

It is also the home of the context graph — the spine of the Data Context Wizard. Every fact it serves carries provenance: which node, which contributing agent, and when it was last updated. When two records conflict — two owners for the same table, two row counts for the same run — it surfaces both and flags them for a human, rather than silently picking one.

  • Natural-language catalog search. search_datasets returns ranked results with relevance scores, filterable by platform, type, tags, and quality score.
  • One-call context. get_context returns schema, lineage, quality, freshness, trust score, and documentation for an asset in a single response — what an agent (or a human) needs before touching a table.
  • Lineage and blast radius. get_lineage traverses upstream and downstream with column-level lineage where available; assess_impact classifies the severity of a change by what it would break — dashboards, models, pipelines.
  • The context graph. query_context_graph combines lineage, decision history, and related incidents for an entity in one contextual answer; traverse_lineage_graph and get_decision_history (with hash-chain proof) drill into each.
  • Governed contributions. contribute_to_graph writes nodes and edges through the governed path — PII scrubbing and authority checks before anything persists, with an audit hash returned.
  • Semantic layer resolution. resolve_metric maps an ambiguous name (“revenue”, “MRR”) to its canonical definition, returning all candidates when several match.
  • Freshness with a straight face. check_freshness scores freshness against an SLA and time-sensitive answers carry an as-of timestamp — a stale entry is flagged as stale, not reported as current.

“Find the authoritative table for customer subscriptions — not the copies.”

“Show me everything downstream of raw.stripe_payments, down to column level.”

“What’s the full context on analytics.mrr_daily — schema, freshness, quality, and who owns it?”

“When someone says ‘active users’, which definition do we actually mean?”

“What decisions have agents made about this table in the last 30 days?”

  • Catalogs — DataHub, OpenMetadata, AWS Glue, Purview, Dataplex, and other metadata sources feed search and lineage.
  • Warehouses and lakehouses — Snowflake, BigQuery, Databricks; plus Iceberg REST catalogs.
  • dbt — models, lineage, and test results.

See the connector catalog. We don’t ship native connectors for Atlan, Alation, or Collibra today — here’s what to do instead.

The agent starts in 🟡 Evaluation on a realistic sample estate — the search ranking, graph traversal, and impact-analysis algorithms are the real thing. It earns 🟢 Connected per system through a passing live test. See Verify your setup.

  • Search runs on keyword and graph signals by default. On-device semantic embeddings are an explicit opt-in; if the embedding model isn’t available, search degrades gracefully to keyword search rather than fusing meaningless vectors into the ranking.
  • The agent can never promote its own writes to authoritative. Marking a source authoritative requires a named human approver — no exceptions.
  • If a query returns nothing, it says so. It will not infer a catalog entry the graph doesn’t contain.
  • Bulk harvest of very large estates is still being hardened — start with a scoped schema, not the whole estate.