Skill: eval-benchmark
Download: eval-benchmark.zip ·
Install per the toolkit page, then run /eval-benchmark in Claude Code.
“Do the agents actually work on our data?” deserves a number, not a vibe. This skill builds that number with you during onboarding — from your own history, graded against your own systems of record — and then turns every failure into an optimization step.
How it works
Section titled “How it works”1 · Collect — mines what your team already asks: the most-repeated queries in your warehouse’s query history, the models and tests that break most in dbt, the questions behind your most-viewed dashboards, and the questions your team knows by heart (“what broke last quarter?” makes the best eval cases). Target: 20–50 cases across insight, SQL, lineage/impact, quality/incident, and cost questions.
2 · Freeze — for each case, the expected answer is captured from your system of
record — reference SQL run against your warehouse, the catalog entry, the dbt graph —
never from an agent’s output. Cases land in a committed eval-set.json (a
template ships in the bundle) with a grading method per
case: exact match, numeric-with-tolerance, set containment, or a rubric judged against
the frozen reference. Time-sensitive cases keep their reference SQL so truth is re-frozen
fresh on every run day.
3 · Benchmark — every case is asked verbatim in a fresh context, graded, and recorded:
pass/fail, the answer given, which agents/tools ran, and a failure class
(wrong-data, missing-context, wrong-tool, refused, error, stale-graph).
Output: a run file plus a scoreboard —
Score: 34/42 (81%) · Prior run: 29/42 (69%)By domain: insights 8/10 · context 9/9 · lineage 6/8 · quality 6/7 · cost 5/84 · Hill-climb — failures are the optimization queue, worked in yield order:
missing-context/stale-graph failures get fixed in the
Data Context Wizard (curate the missing definition,
promote it, re-harvest the stale slice — the cheapest fix, and the clearest demo of the
compounding loop); wrong-data failures get their connection scope — or occasionally the
frozen reference itself — corrected with your data owner; wrong-tool/refused/error
failures are ours: each is filed as a reproducible case through the
feedback pipeline, and our team works them during your onboarding.
Failed cases re-run after each fix batch; the full set re-runs at milestones.
5 · Keep it — the eval set is yours: committed in your repo, extendable (each new incident should become a case), re-run every two weeks during onboarding and after every product upgrade. Day-0 vs. week-2 vs. pilot-end is the evidence pack the pilot decision reads from — and your standing regression suite afterward.
The honesty rules
Section titled “The honesty rules”Expected answers are frozen from your systems, never from agent output. A failed case is a finding, recorded as-is. Every score in a report comes from an actual graded run — the same verify-before-claim rule as everything else here.
Who owns what
Section titled “Who owns what”The eval set and every run file live in your environment. Failure analyses that leave your environment for our product backlog do so with your permission, case by case. What we get out of it — a benchmark per customer that our team hill-climbs against — is exactly what you get out of it: the product measurably improving on your workload, with receipts.