Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

| Plugin | Use When | Tools |
| ------ | -------- | ----- |
| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz` |
| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz` |
| [factory](factory/) | Take a request from scope to a shipped pull request unattended, on Claude Code or Codex. | `run` |
| [bootstrap](bootstrap/) | Prepare any repository for agent work: probe its stack and write its root AGENTS.md from the workflow's SDLC lessons, on Claude Code or Codex. | `agents-md` |

Expand Down
2 changes: 1 addition & 1 deletion dev/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "dev",
"version": "3.5.0",
"version": "3.7.0",
"description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.",
"author": {
"name": "Tobrun"
Expand Down
16 changes: 14 additions & 2 deletions dev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling.

The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The durable context is deliberately small: the code, its tests, the active spec under `.dev/{plan-name}/`, and three repo-tracked registries the skills maintain in the consuming project - `docs/decisions.md` (design decisions with their argued alternatives, read only after a review forms its findings), `docs/contracts.md` (boundary guarantees, read as premises before a review walks the diff), and `docs/dependencies.md` (machine-checkable module dependency rules, enforced by `ship`).
The files under `.dev/{plan-name}/` are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight.
Alongside them, `docs/architecture.md` is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers.
Expand Down Expand Up @@ -32,7 +32,8 @@ Executes a spec's change sets at the layer each `tests:` scenario is tagged with
Enforces outcomes rather than rituals: every scenario gets a real test at its tagged layer, written with the code and run to green, with no red step.
Failing-test-first is reserved for bug fixes, where one red run is the proof the issue was actually reproduced.
Runs independent change sets in parallel as waves of subagents batched by disjoint file lists in spec order, committing each change set and appending to a running `implementation-notes.md` that logs any deviations forced by an edge case.
Each subagent starts from a brief that `scripts/change-set-brief.py` cuts from the spec for its change set and runs only its own tests; the spec's Validation block runs once per wave as the gate before its commits, and e2e, benchmark, and coverage commands run once at the end.
Each subagent starts from a brief that `scripts/change-set-brief.py` cuts from the spec for its change set and runs only its own tests; the spec's Validation block runs once per wave as the gate before its commits, scoped by `scripts/impact-scope.py` to the packages the wave reaches, and e2e, benchmark, and coverage commands run once at the end.
The repository's required pull-request commands run locally only at the branch's impact; what lies outside it is deferred to the pull request's CI, unless a lockfile, root config, toolchain pin, or CI workflow changed, which makes every package impacted.
Once every change set is committed, drives the real app against a mocked environment, loops until every e2e scenario passes, then renders the e2e report: screenshots per scenario for frontend systems, Test Scenario and Data Model State tables for everything else.
Build never pushes or opens a PR; `ship` does, once the change is hardened and reviewed.

Expand All @@ -50,13 +51,22 @@ Claude Code, Codex, and opencode use their native parallel subagent facilities.
Renders `review_N.html` alongside the `review_N.md` file for reviewer handoff.
Phase 3 commits what the gauntlet fixed, pushes, and opens the pull request automatically - a draft when the review verdict is BLOCK - with an Evidence section that shows the change working: screenshots from the e2e run for a system with a frontend, published to a `pr-evidence` branch so they render inline, or labeled before/after state otherwise, and for every bug fix the reproducing test shown red on the merge base and green on the branch.
A deterministic check (`pr-evidence.py check`) gates the PR body, so a PR cannot open on a placeholder or a data URI, and the phase then follows required checks to green.
Local gates in every phase run at the change's impact, so the full merge gate is the PR's own CI: a PR carrying a check deferred to CI opens as a draft and is marked ready once its required checks pass, and a run that opens no PR runs the full set locally.

### commit

Groups all pending changes into granular, logically-separate commits - splitting within a file when needed - with structured messages: a `type(scope):` subject, `What:`/`Why:` body, optional `Considered:`/`Constraint:`/`Directive:`/`Symptoms:` sections, and `Severity:`/`Risk:` metadata trailers.
After committing, syncs drifted docs, captures durable decisions and contracts from the commit bodies into `docs/decisions.md` and `docs/contracts.md`, and runs the architecture checker over the batch so `docs/architecture.md` stays current with every commit, which is what keeps the registries trustworthy without excavating git history later.
Pushes by default; say "commit only" to skip the push.

### reflect

Turns what past runs measured into evidence about the skills themselves, so the next change to the workflow fixes something that actually recurred.
Every `scope`, `scope-review`, `build`, and `ship` run appends a journal entry to `~/.dev-workflow/memory/` (see Run metrics below); `reflect` has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source.
A checker (`scripts/claims.py add`) accepts a batch only when every quote is found verbatim at the transcript line or journal entry it cites, then rebuilds per-skill pages and a ranked list of threads in which every line cites a claim id.
The user picks a thread, or retracts a claim that misreads its evidence (a retracted claim stays on record so the same evidence cannot bring it back), and the skill writes a scope brief with an eval case built from the real run.
It never edits a skill: the fix goes through `scope` and `build` in this repository, and `/dev:reflect resolve {thread} {sha}` records the commit, after which a new claim in that thread reopens it.

### to-pitch

Packages a finished change - its spec, implementation notes, and e2e evidence - into one buy-in document: demo first, then why, what changed, how it was verified, and how to try it.
Expand All @@ -71,6 +81,8 @@ Cannot enforce a merge gate, so it says so plainly and produces an honest pass/f
`scope`, `scope-review`, `build`, and `ship` each start by snapshotting the run with `scripts/skill-metrics.py start` and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output.
The numbers come from the session transcript under `$CLAUDE_CONFIG_DIR` (default `~/.claude`) and from git, never from the model's recollection.
Every run appends a row to `.dev/metrics.jsonl` in the consuming repository, and the table compares the run against the median of earlier runs of the same skill, which is where a skill's cost and catch rate become visible over time.
The same call appends an entry to the cross-repository run journal under `~/.dev-workflow/memory/journal/` (or `$DEV_MEMORY_DIR`): the friction signals measured from the transcript (interrupts, denied and failed tool calls, the user's own turns, and how often each deterministic checker ran and failed), each with its transcript line, plus at most three `--friction` lines in which the skill names where the run fought its own instructions.
That journal is what `reflect` consolidates. It stays on the machine; set `DEV_MEMORY_DIR=off`, or `"memory": {"enabled": false}` in `.dev/config.json`, to keep a repository out of it.

## Jira integration

Expand Down
6 changes: 3 additions & 3 deletions dev/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,12 +4,12 @@ Status: maintained

Eval definitions for the `dev` plugin's skills: realistic prompts and objective assertions used to check whether a skill change preserved behavior.

- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 7 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz`.
- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 8 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz`.
- `results.md` - the record of the most recent full run: scores, methodology, and findings.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`, `check-tests.py`, `pr-evidence.py`), run by `scripts/validate.sh` as check D01.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`, `check-tests.py`, `pr-evidence.py`, `impact-scope.py`, and the memory scripts `skill-metrics.py` and `claims.py`), run by `scripts/validate.sh` as check D01.

`build` runs in `"functional"` mode (a real fixture, a real subagent run, assertions checked against the actual output).
`scope`, `scope-review`, `commit`, `ship`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.
`scope`, `scope-review`, `commit`, `ship`, `reflect`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.

## When to run these

Expand Down
21 changes: 16 additions & 5 deletions dev/evals/build.json
Original file line number Diff line number Diff line change
Expand Up @@ -80,17 +80,28 @@
{
"id": "ci-parity-before-pr",
"prompt": "Implement this spec.",
"fixture": "a repo whose spec Validation block runs unit/lint/build, while .github/workflows/pr.yml also requires a project-owned screenshot-matrix command. The active-month screenshot fails deterministically on the first UTC day, although the feature's tagged e2e scenarios pass.",
"fixture": "a repo whose spec Validation block runs unit/lint/build, while .github/workflows/pr.yml also requires a project-owned screenshot-matrix command over the UI package. The spec changes that UI package, and the active-month screenshot fails deterministically on the first UTC day, although the feature's tagged e2e scenarios pass.",
"assertions": [
"The pull-request workflow is inspected and the screenshot-matrix command is run before the user is offered a PR",
"The pull-request workflow is inspected and impact-scope.py is run before the required commands, and the screenshot-matrix command runs before the user is offered a PR because the change reaches its inputs",
"A required command that already ran green on the unchanged final tree is recorded, not run a second time",
"The screenshot failure is diagnosed and fixed rather than labeled pre-existing, flaky, unrelated, or an accepted deviation",
"The screenshot failure is diagnosed and fixed rather than labeled pre-existing, flaky, unrelated, deferred to CI, or an accepted deviation",
"The fix removes the time-dependent assumption without adding a retry, sleep, timeout increase, or looser assertion",
"The full screenshot-matrix command is rerun and green before the PR question",
"The discovered CI commands and outcomes are recorded in implementation-notes.md, each as its command finishes",
"The screenshot-matrix command is rerun and green before the PR question",
"The discovered CI commands, the scope each ran at, and their outcomes are recorded in implementation-notes.md, each as its command finishes",
"If the user approves a PR, required checks are watched to a terminal state and a deterministic failure is fixed and pushed rather than merely reported as restarted"
]
},
{
"id": "impact-scoped-gates",
"prompt": "Implement this spec.",
"fixture": "an npm-workspaces repo with packages/api and packages/web, each with its own test script, and a pr.yml that runs every workspace's tests; the spec's two change sets edit only packages/api",
"assertions": [
"Each wave gate runs impact-scope.py --base HEAD and runs the Validation commands over packages/api (and any workspace that depends on it), never the packages/web suite",
"The final CI-parity run is scoped by impact-scope.py against the merge base, and the packages/web tests are recorded in implementation-notes.md as deferred to CI: outside impact rather than run or silently dropped",
"No required command whose inputs the change reaches is deferred because it is slow",
"The closing message points at ship, whose PR's CI runs the deferred commands"
]
},
{
"id": "jira-disabled",
"prompt": "Implement this spec.",
Expand Down
Loading
Loading