Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
79 changes: 79 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -381,6 +381,85 @@ mb transform-tag delete 5 --yes
| ------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `--yes` | Skip the interactive confirmation prompt. In non-TTY contexts the prompt is skipped automatically (kubectl/gh/docker convention). |

## Transform tests

CRUD and run on `/api/ee/transform-test`. Requires Metabase v65 or newer with the `transforms-testing` premium feature. A transform test pins a transform's behaviour: every table the transform reads is replaced by an `input` fixture, the transform runs into a temp table, and each `expectation` checks that output. The temp tables are dropped when the run ends. The transform and the expectations read only those temp tables; a `format: "sql"` input runs verbatim against the transform's source database, so it may read real tables. Only query transforms (native SQL or MBQL) can be tested, on Postgres, MySQL, H2, ClickHouse, Redshift or SQL Server.

An input names its `table` and carries either `format: "sql"` with a `sql` query or `format: "rows"` with `columns` (each a `name` and a `cast_type` the warehouse accepts as a `CAST` target) and `rows`; every row carries exactly the declared columns.

An expectation is either `type: "empty"` with the `sql` that must return no rows, or `type: "equals"` with `format: "rows"` and the `columns` and `rows` the output must hold exactly. Expectation names are unique within a test. Expectation SQL may name only the transform's target table and its declared input tables, which are rewritten to the run's temp tables; any other table, or a column qualified by a table name rather than an alias, is refused with `transform-test.unremapped-reference`.

Create and update bodies are closed at every level, so a test read back with `get --full` has to shed `id`, `entity_id`, `creator_id`, `created_at` and `updated_at` before it can be sent back:

```sh
mb transform-test get 1 --full --json \
| jq 'del(.id, .entity_id, .creator_id, .created_at, .updated_at)' \
| mb transform-test update 1
```

### `mb transform-test list`

```sh
mb transform-test list --json
mb transform-test list --transform-id 1
```

| Flag | Description |
| --------------------- | ------------------------------------ |
| `--transform-id <id>` | Only the tests of this transform id. |

### `mb transform-test get <id>`

The compact form carries the id, transform, name and description; `--full` adds the `inputs` and `expectations` themselves.

```sh
mb transform-test get 1
mb transform-test get 1 --full --json
```

### `mb transform-test create`

```sh
mb transform-test create --file transform-test.json
```

| Flag | Description |
| --------------- | ----------------------- |
| `--body <json>` | Inline JSON body. |
| `--file <path>` | Path to JSON body file. |

### `mb transform-test update <id>`

Only the fields the body carries are patched; `inputs` and `expectations` replace what is stored rather than merging into it.

```sh
mb transform-test update 1 --body '{"name":"renamed"}'
```

### `mb transform-test delete <id>`

```sh
mb transform-test delete 1 --yes
```

| Flag | Description |
| ------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `--yes` | Skip the interactive confirmation prompt. In non-TTY contexts the prompt is skipped automatically (kubectl/gh/docker convention). |

### `mb transform-test run <id>`

Runs the transform against the fixtures and reports what each expectation found. Exits 1 when the test does not pass, so it drops straight into CI.

```sh
mb transform-test run 1
mb transform-test run 1 --json
mb transform-test run 1 --timeout 300000
```

| Flag | Description |
| ---------------- | -------------------------------------------------------------------------- |
| `--timeout <ms>` | Request timeout in ms (default 30000); the run is one synchronous request. |

## Databases

Read warehouse metadata from `/api/database`. The `db` group exposes the full database list, the per-database record with optional table/field hydration, schema and table inspection, and the two manual-sync triggers.
Expand Down
2 changes: 1 addition & 1 deletion packages/cli/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@metabase/cli",
"version": "0.4.0",
"version": "0.4.1",
"description": "Metabase CLI",
"license": "AGPL-3.0",
"repository": {
Expand Down
5 changes: 4 additions & 1 deletion packages/cli/skill-data/core/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Top-level command groups (run `mb <group> --help` to discover verbs):

```
auth | db | table | field | upload | content-translation | query | card | dashboard | snippet | segment | measure | collection | library
document | glossary | timeline | timeline-event | transform | transform-job | transform-tag | alert | subscription | setting
document | glossary | timeline | timeline-event | transform | transform-job | transform-tag | transform-test | alert | subscription | setting
search | dependency | git-sync | setup | eid | uuid | upgrade | skills
```

Expand Down Expand Up @@ -171,6 +171,9 @@ This file is enough for any single-command task. For anything deeper, load the r
- **`notification`** — scheduled delivery: question alerts (`mb alert`) and dashboard subscriptions (`mb subscription`). Choosing between them, the two schedule/recipient contracts, channel prerequisites, testing a send.
<!-- requires: transforms -->
- **`transform`** — transform body JSON, create + run-with-wait, run inspection, tags, jobs.
<!-- /requires -->
<!-- requires: transformTests -->
- **`transform-test-plan`** — planning transform tests.
<!-- /requires -->
- **`document`** — Metabase documents (TipTap body, embedding cards).
<!-- requires: remoteSync -->
Expand Down
82 changes: 82 additions & 0 deletions packages/cli/skill-data/transform-test-plan/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
---
name: transform-test-plan
description: Derive a comprehensive test plan for a transform — the fixture cast, the expectations, hand-derived expected rows, and a coverage matrix — from the model's declared design. Covers input partitioning (zero-case, multiplicity, dirty rows), grain / conservation / recomputation / conformance checks, and known-quirk conventions. Load when the user wants tests planned or written for transforms — "write tests for my transforms", "is my model right", "test plan for this pipeline", "add data quality checks" — whether the model is mid-build or already deployed. The `mb transform-test` command and body shapes live in the `transform` skill; this one decides what to test.
allowed-tools: Read, Write, Edit, Bash, AskUserQuestion
requires: [transformTests]
---

# Planning transform tests

Turn a transform into a fixture cast that proves the logic on small known rows, expectations that state the model's invariants, and a coverage matrix showing what's checked and what's deliberately not. Mechanics — the `inputs`/`expectations` body shape and every `mb transform-test` verb — live in the `transform` skill (`mb skills get transform`); load it before authoring, and never restate it here. The two references this skill routes through, `references/checklist.md` and `references/checks.md`, come with `mb skills get transform-test-plan --full`, or read them from `mb skills path transform-test-plan`.

Every check derives from what the model **declares** — detected from its SQL, confirmed with its owner — never from conformance to a modeling doctrine. One plan serves two moments: while the model is **built**, checks pin each design decision; once **deployed**, the same SQL screens production tables for anomalies.

## Operating rules

- **Detect, then derive.** Classify what the transform is and which conventions it uses (the checklist); derive checks only from that. Star, one-big-table, partial denormalization — all fine; never flag a style.
- **Judgment calls go through the checklist.** The session's autonomy setting governs which answers you supply yourself and which you bring to the user — it never makes the checklist a formality. Every answer you supply yourself is recorded in the plan as a stated assumption, paired with the named expectation that enforces it — reversing the decision then breaks a test, not a paragraph. And regardless of setting, when genuinely unsure, ask — a wrong-but-confident grain poisons every downstream check.
- **Expected rows are derived by hand** from the fixture story and business meaning — never captured from the transform's output, which asserts only that the transform equals itself.
- **Batteries stay off until declared structure switches them on.** No snapshot-density checks without a snapshot, no version-history checks without effective/end/current columns. An empty section beats a speculative one.

## The procedure

1. **Profile the real inputs**: per table, row count; per column, min/max/distinct-count/null incidence; orphan counts across declared links (`mb field summary`, `mb query`). Profiling feeds domains, bounds, null partitions, and key candidates — and every fixture edge cites the real-data condition that warrants it, with its count ("the warehouse has 67 ship-before-order rows").
2. **Detect.** Read the transform's SQL (`mb transform get <id> --full --json`) for: grain candidates (`GROUP BY` keys, the joins' driving table), join types (orphan handling), correlated aggregates (stored aggregates), `<entity>_<attr>` naming (copies), effective/end/current columns (version history), `now()`/`current_date` (volatile columns). A transform whose grain you cannot state in one phrase is itself a finding — raise it before writing any test.
3. **Confirm.** Walk [references/checklist.md](references/checklist.md). Three stages: classify the model, per-table declarations, per-column declarations. Each question carries its detection hint; answer autonomously where the hint resolves, ask where it doesn't. In build-along mode these are design questions — treat an undecided answer as a decision to make together, not a blocker.
4. **Derive.** Route every output column through [references/checks.md](references/checks.md) — declared property → expectation shape, fixture implication, expected-row convention. Read it in full once per plan; it is the plan's content.
5. **Design the fixture cast**: one small cast per transform (≈5–10 rows per table), human-named rows ("Alice Premium"), every row a named edge — zero-case entities for every outer join and aggregation, ≥2-member groups for every grouping and join, one dirty row per screenable defect, boundary dates. Document it as a table (row → attributes → purpose) in the plan.

Express the cast as literal rows rather than as a query, so the story stays legible in the test itself; a query is worth it only when the rows are mechanical to generate. A cast is written against one warehouse and does not carry to another.

6. **Hand-derive the expected rows**, arithmetic recorded in the plan (premium: 3 orders / 350.50 / 3.0). Every fixture row's fate appears in some expected cell. Pin NULL-vs-0-vs-empty for every zero-case row — that cell is the null policy's only enforcement.
7. **Author the test.** One test per transform, per coherent story: an `equals` pinning the output, declaring every output column, since `equals` compares only the columns it declares, and `empty` expectations stating the invariants that survive a change to the cast.

Give each expectation a name that states the invariant, because the name is what a failure leads with ("revenue never negative", not "check 3"), and open each `empty` expectation's SQL with a `--` contract comment: the invariant, and the failure modes it catches ("catches both dropped orders and join fan-out").

Where one invariant applies to several transforms, duplicate the SQL. Nothing is shared between tests, so give each copy its own name and comment rather than one that only makes sense next to its twin.

8. **Emit the coverage matrix** in the plan: rows = output columns with the table's grain; columns = check classes (grain, conservation, recomputation, conformance, domains, referential integrity, temporal, screens); cells name the covering expectation or expected-row cell, or state `gap: <reason>`. Scan shared-attribute columns across transforms for cross-table agreement obligations. Empty cells are honest; silent gaps are not.
9. **Prove the tests have teeth.** Once per test: corrupt one expected cell, `run`, confirm `cell-mismatches` names exactly that column; revert. Then perturb one input cell and confirm exactly the declared output cells move. An `empty` expectation that only selects output rows passes against an empty output, so a green test can still be vacuous.

## Severity: error, or the tolerated oddity

Every expectation is pass or fail, and one failure fails the run. **error** = forbidden, and it becomes an `empty` expectation. There is no warn severity, so **never author an expectation you expect to fire**: a permanently red test trains everyone to ignore the result, and a red run gates everyone else's work.

A tolerated-but-surfaced oddity — one the owner lives with, like orphan rows or ship-before-order dates — gets encoded the two ways that hold: **pin it in the `equals` rows** (an orphan passing through with NULLs is a cell in the expected output, so reversing the tolerance breaks the test), and **record it in the plan's known-quirks list** with its real-data count, so the tolerance stays a conscious choice.

Under budget pressure cut business-rule checks first, then cross-table structure checks; never single-column screens (domains, ranges, nulls) — cheapest, and the last line.

## When a check exposes a live bug

Non-negotiable: **never soften the test to green** — expected values state correct behavior; matching them to buggy output documents the bug as intended — and **surface the finding with its blast radius at both scales**, fixture ("1530.24 of 1600.74 fixture dollars survive") and warehouse ("908 of 2,050 orders dropped"). What happens next follows the session's terms, not a fixed protocol: propose and apply the fix now (when the user wants it or the autonomy setting covers it), or — when the fix must wait — hold the correct expectation and record the red in the plan with the minimal fix body. A stored test has no red-by-design state, so a deferred fix must be visible in the plan or the test reads as broken. Either way, once green the test stays as the regression guard.

## When the doctrine doesn't apply

The vocabulary follows Kimball's dimensional modeling (grain, additivity, conformed attributes, slowly changing dimensions) — precise, widely understood terms. Real models are Kimball-inspired at most; no check may score adherence:

- Full calendar date dimensions are rare. Never demand one; test date _semantics_ — ranges, orderings, volatile derivations.
- Surrogate keys are doctrine, natural-key joins are practice. Test whichever key the model declares; never flag natural-key joins.
- "No NULL FKs / no NULL attributes" is doctrine routinely dropped. Null policy is three independent declarations — measures, attributes, FKs — each detected and confirmed, never presumed.
- One-big-table is legitimate: it still has a grain, its copies still need agreement checks, its functional dependencies still hold.
- A transform-level `ORDER BY` has no testable effect — output tables carry no row order and the `equals` comparison is a multiset. Flag it as probable dead weight (clustering hints aside); never write an ordering expectation.

## One transform at a time

Each test covers one transform. When the transform under test reads another transform's output, that target table is an input like any other: declare it and fake it. Derive those rows from the base transform's own expected output, so the two tests tell one story, and note the coupling in both plans — changing the base's expected rows means changing this test's input.

## Worked example, condensed

`orders` + `customers` → _enriched_orders_ (order grain; LEFT JOIN attaches `customer_name`, `tier`).

Cast, 7 orders: two tiers; Alice and Carol with 2 orders each (multiplicity); two never shipped (zero-case for the shipping join); order 106 shipped before ordered (real oddity, count cited); order 107's customer_id matches no customer (orphan). Two `rows` inputs, one per source table.

Derived, per catalog: grain uniqueness on `order_id`; row and amount conservation from input to output, as scalar subqueries over the target and the seeded input; `tier` domain ⊆ {standard, premium}; orphan = keep-with-NULLs → an expected row pinning the NULL pass-through; ship-before-order tolerated → the plan's known-quirks list with the warehouse count. The whole output, every column, pinned in one `equals`, hand-computed, arithmetic in the plan. The inner-vs-LEFT-join bug this cast catches — unshipped orders silently dropped from revenue — is what the zero-case rows exist for.

## Don't

- Don't author expectations before loading the `transform` skill — the body shape, the closed create/update contract, and the verb flags live there.
- Don't capture expected rows from the transform's own output — hand-derive them or they assert nothing.
- Don't emit checks for structure the model doesn't declare (snapshot density, version history, bridge weights) — an inapplicable battery buries real findings.
- Don't let a fixture cast go all-clean — no zero-case, no orphan, no dirty row proves the happy path and nothing else; the bugs live in the edges.
- Don't write an expectation you expect to fail.
- Don't surface bare check-ids ("per C3…") to the user — name the check in plain words; the ids are for your cross-referencing, not their reading.
Loading
Loading