Skill: onboard-data-workers
Download: onboard-data-workers.zip ·
Install per the toolkit page, then run /onboard-data-workers in Claude Code.
This is the skill your onboarding engineer runs at kickoff — and that your team keeps for expansions (new schemas, new teams, new workspaces). It drives the whole journey and won’t mark a phase complete without its checkpoint passing live, in the session.
The eight phases
Section titled “The eight phases”| Phase | What happens | Checkpoint |
|---|---|---|
| 0 · Intake | Captures the eight facts everything depends on: stack, coding agents, the first schema that matters, npm proxying, token, model key. | All eight answered; token in hand |
| 1 · Install | Installs the agent swarm into your coding agent (claude mcp add data-workers -- npx -y dw-claw, or the Cursor/Codex/OpenCode equivalent) and smoke-tests on sample data — before any credential exists. | A sensible sample-data answer in your own client |
| 2 · Connect | Your team creates least-privilege read credentials and sets each system’s environment variables. Credentials never leave your machines. | Env vars set, client restarted |
| 3 · Verify | Runs a real connection test per system and records the honest three-state result. Nothing proceeds until every in-scope system is 🟢. | Every system 🟢 from a live test |
| 4 · Warm the graph | Harvests your scoped schema into the Data Context Wizard, spot-checks counts with your data owner, curates the ~20 facts that matter, and promotes them with a named approver. | Scoped schema harvested + first facts authoritative |
| 5 · Console | Stands up the Spellbook Data Catalog (research preview, local), registers and tests connectors, and sets your Conductor mode — Approve-edits first, always. | Console live, connectors tested, mode recorded |
| 6 · Eval baseline | Runs the eval-benchmark skill to freeze your eval set and record the day-0 score — the number the rest of onboarding is measured against. | Eval set committed, baseline scored |
| 7 · Report | Writes the onboarding report: per-system states with how each was verified, graph status, console/Conductor setup, eval baseline, and open items — each filed through the feedback pipeline with a reference. | Report delivered, check-in booked |
The rules the skill enforces
Section titled “The rules the skill enforces”- Verify before “Connected” — a credential set is not a connection; only a passing live test flips a system 🟢, same as everywhere else in the product.
- Least privilege, customer-held credentials — read-only scoped service accounts, values never pasted into chat or written into reports.
- Honest maturity — preview surfaces are named as preview; unsupported operations get filed, not faked.
- Scoped first — one schema, a handful of engineers, then expand on green.
What you’re left with
Section titled “What you’re left with”A working, verified deployment; a warmed scoped graph with named approvers; a running console; a day-0 benchmark score; and a report where every line is something that happened in the session — the artifact your security team, your data lead, and our account team all read from.