Skip to content

Latest commit

 

History

History
174 lines (133 loc) · 19.1 KB

File metadata and controls

174 lines (133 loc) · 19.1 KB

dev

Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling.

The skills chain loosely rather than as a rigid pipeline: /scope interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; /scope-review is an optional step for a large or complex change, never a gate in front of build: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands build a spec ready to implement; /build executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; /ship runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; /scope-quick and /ship-quick are the short loop for a change that is already small - most often the fixes a ship review asked for - writing minimal change sets with no interview, and re-shipping with one reviewer instead of the gauntlet and panel; /commit groups pending changes into granular commits with structured what/why messages; /reflect consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to scope as a brief; /to-pitch and /to-quiz turn finished work into a buy-in doc or a comprehension check. The durable context is deliberately small: the code, its tests, the active spec under .dev/{plan-name}/, and three repo-tracked registries the skills maintain in the consuming project - docs/decisions.md (design decisions with their argued alternatives, read only after a review forms its findings), docs/contracts.md (boundary guarantees, read as premises before a review walks the diff), and docs/dependencies.md (machine-checkable module dependency rules, enforced by ship). The files under .dev/{plan-name}/ are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight. Alongside them, docs/architecture.md is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers. Every producing skill renders its own output as self-contained HTML under /tmp/{project-slug}/reports/. It publishes only when the user requests a shareable link and the host provides an artifact-publishing tool. Every skill is explicit-invocation only: Claude Code and Pi use disable-model-invocation: true, the generated Codex distribution uses agents/openai.yaml with allow_implicit_invocation: false, and opencode enforces it with a permission.skill rule set to ask (see opencode installation below). Skills recommend the next step rather than launching each other.

Skills

scope

Specs a change by interviewing for the real problem behind the request, cataloging every design decision (with a subagent blind-spot pass on full-size changes), and arguing each one against alternatives in a ✓/✗/?/⚠/⊘ notation with evidence marks. Writes a self-contained spec at .dev/{plan-name}/spec.md - research decisions, scope with invariants and a Validation block of the repo's real commands, and a change plan of numbered change sets each ending in a layer-tagged tests: line - designed as a fresh-context handoff to build. A checker (scripts/lint-spec.py) enforces the spec's mechanics - unique slugs, argued alternatives, echoes that match their decision, tagged test scenarios, at most 25 scenarios per change set - so the prose stays about judgment. On a clean spec it prints the build waves the file lists allow and the shared files that make change sets wait, so the plan is shaped for parallel work before build starts. Promotes durable decisions to docs/decisions.md and cross-boundary invariants to docs/contracts.md, renders an expandable-card spec view, and has a reverse mode that audits the implicit decisions already embedded in existing code.

scope-quick

The minimal scope: turns the Blockers of the latest ship review, or a small request, into change sets in .dev/{plan-name}/spec.md that build can execute straight away. It runs no interview, argues no decisions, launches no subagents, and skips the ledger work, the HTML render, and the retrospective; a decision is recorded only where the code offered a real choice. Review fixes are appended to the existing change plan with continued numbering, each carrying the review's triggering scenario as its test, and the same lint-spec.py loop keeps the result mechanically sound. When a finding needs a recorded decision flipped or user-visible scope changed, it stops and points at scope instead of guessing.

scope-review

Optional: build runs on any spec scope finished, and scope advises this review only when the change is large or complex. Reviews a settled spec with agents that did not write it and refines it in place, before build starts - the cheapest review in the chain, since a defect caught here costs a spec edit instead of a re-implementation, and the loop makes those edits itself. Gates on scope's own lint-spec.py first (a mechanically unsettled spec is sent back, not reviewed), then loops up to two rounds of panel, verification, and refinement: four lenses in parallel - feasibility (does the plan survive contact with the repo: named files, assumed hooks, claimed prior art, stated premises about current behavior), completeness (failure paths, second-order work, and affected call sites the spec never mentions), consistency (change sets versus their linked decisions and the project's ledgers, semantically), and testability (will the tests: lines produce real proof at their tagged layers) - with every BLOCK and CONCERN adversarially verified on ship's transport and aggregation machinery before a refine agent may touch the spec. Refinements follow a strict authority order (recorded user intent beats the repo's reality beats settled decisions beats spec prose) and edit spec.md only; anything that would flip a decision, change user-visible scope, or add a dependency is never auto-applied - the run ends by asking the user those as decisions, one at a time with alternatives and tradeoffs, and applies the answers to the spec before finishing. An APPROVED verdict means every finding was refined or answered: build can start directly. Only an answer that invalidates the premise or opens a genuinely new effort defers to scope, recorded in .dev/{plan-name}/spec-review_N.md with its open question.

build

Executes a spec's change sets at the layer each tests: scenario is tagged with - unit for business logic, integration for real cross-component seams, e2e for driving the actual running application. Enforces outcomes rather than rituals: every scenario gets a real test at its tagged layer, written with the code and run to green, with no red step. Failing-test-first is reserved for bug fixes, where one red run is the proof the issue was actually reproduced. Runs independent change sets in parallel as waves of subagents batched by disjoint file lists in spec order, committing each change set and appending to a running implementation-notes.md that logs any deviations forced by an edge case. Each subagent starts from a brief that scripts/change-set-brief.py cuts from the spec for its change set and runs only its own tests; the spec's Validation block runs once per wave as the gate before its commits, scoped by scripts/impact-scope.py to the packages the wave reaches, and e2e, benchmark, and coverage commands run once at the end. The repository's required pull-request commands run locally only at the branch's impact; what lies outside it is deferred to the pull request's CI, unless a lockfile, root config, toolchain pin, or CI workflow changed, which makes every package impacted. Once every change set is committed, drives the real app against a mocked environment, loops until every e2e scenario passes, then renders the e2e report: screenshots per scenario for frontend systems, Test Scenario and Data Model State tables for everything else. Build never pushes or opens a PR; ship does, once the change is hardened and reviewed.

ship

Runs the quality pass that finishes a change, in three phases; by default all run, and the first two can be requested alone ("gauntlet only", "review only"), or the flow can stop short of the PR ("no PR"). Phase 1 puts the change through deterministic tools that cannot be argued with, looping fresh-context fix agents until every check passes: every linter, type checker, and format checker the repo already configures, a security scan (secrets, vulnerable dependencies, SAST), dead code and duplication introduced by the diff, module dependency rules from docs/dependencies.md, coverage-weighted cyclomatic complexity per function, flakiness runs over diff-touched tests, and mutation testing over the in-scope files. It acquires tools up an explicit ladder - the repo's own tooling, the ecosystem's established tool, or a small repo-fitted script committed under tools/harden/ for reuse - and treats thresholds as recorded decisions in docs/decisions.md, never silently adjusted config. When the gauntlet's fixes touched code, phase 1 ends by re-running the spec's [e2e] scenarios and overwriting the e2e report, so the evidence phase 2 audits describes the post-fix code. Phase 2 reviews the post-fix diff with a read-only panel of concern-focused agents, selected per diff except the always-on simplify lens; the gauntlet checks mechanics, the panel judges meaning. Every non-trivial finding is adversarially verified against the repo. A BLOCK verdict is not yet a human call: ship runs up to two autonomous remediation rounds, each dispatching fresh-context fix agents at the confirmed blockers, re-hardening the touched files, and re-reviewing at the next index; only a blocker that survives both rounds, or that an agent escalates as needing a spec or decision change, is presented to the user. Reads docs/contracts.md boundary guarantees as premises before the panel runs, checks spec conformance - including whether [e2e]-tagged scenarios have a passing entry in the e2e report and whether logged deviations still satisfy the spec - and reconciles verified findings against docs/decisions.md only after judgment, reporting still-holds/reopened/diverged instead of re-litigating settled questions. Claude Code, Codex, and opencode use their native parallel subagent facilities. Pi preserves the same independent two-batch panel by launching isolated pi --print subprocesses with the current provider, model, and reasoning level. Renders review_N.html alongside the review_N.md file for reviewer handoff. Phase 3 commits what the gauntlet fixed, pushes, and opens the pull request automatically - a draft when the review verdict is BLOCK - with an Evidence section that shows the change working: screenshots from the e2e run for a system with a frontend, published to a pr-evidence branch so they render inline, or labeled before/after state otherwise, and for every bug fix the reproducing test shown red on the merge base and green on the branch. A deterministic check (pr-evidence.py check) gates the PR body, so a PR cannot open on a placeholder or a data URI, and the phase then follows required checks to green. Local gates in every phase run at the change's impact, so the full merge gate is the PR's own CI: a PR carrying a check deferred to CI opens as a draft and is marked ready once its required checks pass, and a run that opens no PR runs the full set locally.

ship-quick

The minimal ship, for re-shipping after the fixes a review asked for: one validation run at the diff's impact, one read-only reviewer, and a pull request update followed to green. The reviewer reports each finding of the previous review_N.md as fixed or still open and reads only the diff since that review's recorded head for new blockers; it raises no concerns, nits, or simplifications. There is no gauntlet, no lens panel, no remediation loop, and no HTML report: a standing blocker ends the run with a pointer back to scope-quick. An existing pull request keeps the Evidence and Quality sections the full ship run wrote, with only its review line and open calls updated. Run the full ship for a first ship, or once the change has grown beyond the findings it set out to fix.

commit

Groups all pending changes into granular, logically-separate commits - splitting within a file when needed - with structured messages: a type(scope): subject, What:/Why: body, optional Considered:/Constraint:/Directive:/Symptoms: sections, and Severity:/Risk: metadata trailers. After committing, syncs drifted docs, captures durable decisions and contracts from the commit bodies into docs/decisions.md and docs/contracts.md, and runs the architecture checker over the batch so docs/architecture.md stays current with every commit, which is what keeps the registries trustworthy without excavating git history later. Pushes by default; say "commit only" to skip the push.

reflect

Turns what past runs measured into evidence about the skills themselves, so the next change to the workflow fixes something that actually recurred. Every scope, scope-review, build, and ship run, and every run of their quick variants, appends a journal entry to ~/.dev-workflow/memory/ (see Run metrics below); reflect has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source. A checker (scripts/claims.py add) accepts a batch only when every quote is found verbatim at the transcript line or journal entry it cites, then rebuilds per-skill pages and a ranked list of threads in which every line cites a claim id. The user picks a thread, or retracts a claim that misreads its evidence (a retracted claim stays on record so the same evidence cannot bring it back), and the skill writes a scope brief with an eval case built from the real run. It never edits a skill: the fix goes through scope and build in this repository, and /dev:reflect resolve {thread} {sha} records the commit, after which a new claim in that thread reopens it.

to-pitch

Packages a finished change - its spec, implementation notes, and e2e evidence - into one buy-in document: demo first, then why, what changed, how it was verified, and how to try it.

to-quiz

Renders a context/intuition/what-was-done report with a graded comprehension quiz. Cannot enforce a merge gate, so it says so plainly and produces an honest pass/fail check instead.

Run metrics

scope, scope-review, build, ship, scope-quick, and ship-quick each start by snapshotting the run with scripts/skill-metrics.py start and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output. The numbers come from the session transcript under $CLAUDE_CONFIG_DIR (default ~/.claude) and from git, never from the model's recollection. Every run appends a row to .dev/metrics.jsonl in the consuming repository, and the table compares the run against the median of earlier runs of the same skill, which is where a skill's cost and catch rate become visible over time. The same call appends an entry to the cross-repository run journal under ~/.dev-workflow/memory/journal/ (or $DEV_MEMORY_DIR): the friction signals measured from the transcript (interrupts, denied and failed tool calls, the user's own turns, and how often each deterministic checker ran and failed), each with its transcript line, plus at most three --friction lines in which the skill names where the run fought its own instructions. That journal is what reflect consolidates. It stays on the machine; set DEV_MEMORY_DIR=off, or "memory": {"enabled": false} in .dev/config.json, to keep a repository out of it.

Claude Code installation

/plugin marketplace add tobrun/workflow
/plugin install dev@nurbot

Invoke skills as /scope, /build, and so on.

Codex installation

codex plugin marketplace add tobrun/workflow
codex plugin add dev@nurbot

Invoke skills as $dev:scope, $dev:build, and so on.

opencode installation

opencode reads Claude-format SKILL.md files natively, so no generated distribution is needed. Clone this repository and symlink the source skills into opencode's global skill directory:

git clone https://github.com/tobrun/workflow ~/ws/workflow
mkdir -p ~/.config/opencode/skills
for skill in ~/ws/workflow/dev/skills/*/; do
  ln -sfn "$skill" ~/.config/opencode/skills/"$(basename "$skill")"
done
ln -sfn ~/ws/workflow/dev/references ~/.config/opencode/references

The last symlink keeps the shared references (decision-ledger.md, contracts.md, architecture.md) reachable through the ../../references/ links inside the skills. Symlinks mean a git pull updates the skills in place; restart opencode afterwards, since skills load at startup.

opencode has no disable-model-invocation field (it is ignored harmlessly); skills load through a model-invoked skill tool. Preserve the explicit-invocation policy with a permission rule in ~/.config/opencode/opencode.json:

{
  "permission": {
    "skill": {
      "*": "allow",
      "scope": "ask",
      "scope-quick": "ask",
      "scope-review": "ask",
      "commit": "ask",
      "build": "ask",
      "ship": "ask",
      "ship-quick": "ask",
      "reflect": "ask",
      "to-pitch": "ask",
      "to-quiz": "ask"
    }
  }
}

Invoke a skill by asking for it by name, for example "run the scope skill on this request"; the permission rule makes opencode confirm before loading one. Verify discovery with opencode debug skill. ship runs its review panel through opencode's native task subagents.

Pi installation

pi install git:github.com/tobrun/workflow

Invoke skills as /skill:scope, /skill:build, and so on. Pi loads the source skills directly. A ship review phase starts multiple model processes, one per selected lens and then one per non-trivial finding verifier, so its model usage scales with the panel size.