Skip to content

Research digest: spec-kit, BMAD, OpenSpec, MADR, log4brains — 4 integration candidates, 3 settled non-candidates #38

Description

@minerva-sky

Research loop digest (first run for this repo). Method: feature-level survey of the five nearest neighbors — GitHub spec-kit, BMAD-METHOD, OpenSpec, MADR, log4brains — against what this framework ships today (7 skills, 8 reviewer personas, MCP server, plugin distribution). Every candidate below is fenced against WORLD.md before it appears; things that fail the fence are listed as non-candidates so the next research pass doesn't re-raise them.

The landscape

github/spec-kit (~111k stars, MIT) — spec → plan → tasks → implement workflow, 30+ assistants supported. Two things it does that we don't: (1) a constitution — project principles captured once and referenced in every subsequent phase, not just consulted when someone remembers; (2) single-source generated per-assistant command files, so 30+ integrations don't drift. What it got wrong from our vantage: it's a code-generation pipeline (our explicit anti-goal), and it churns — v0.10.0 removed its --ai flag family outright, breaking every tutorial written before June 2026. Their churn is our argument for markdown-first.

BMAD-METHOD (~49k stars) — the nearest neighbor on personas. V6 ships 12+ personas with behavioral system prompts and handoff protocols, plus scale-adaptive planning depth (a bug fix gets a shallow pass, an enterprise system a deep one). What it got wrong: it swallowed agile project management whole — sprints, story files, scrum-master personas — which is our anti-goal verbatim, and its token burn is a recurring community complaint. The persona-behavior gap it exposes in our members.yml is already filed as #35, and its collaborative rounds as #36 — not re-raised here.

OpenSpec (Fission-AI) — lightweight proposal-first SDD: each change is a folder (proposal / specs / design / tasks), no phase gates, no API keys. Its distinctive idea is the delta model: specs are living current-state truth, and change proposals patch them, with an archive step after implementation. That's a genuinely different memory model from append-only ADRs.

MADR 4.x — the de-facto markdown ADR standard. Our adr-template.md already carries MADR-style Decision Drivers. Two things MADR 4.0 has that we lack: a Confirmation field (how will we verify the implementation actually follows this decision?) and YAML frontmatter for machine-readable status/date/deciders.

log4brains (~1.4k stars, officially "low maintenance mode") — ADR CLI + static-site knowledge base. Its trajectory is the cautionary tale: the runtime app decayed while the markdown files it managed remain useful. That's our Constraints section vindicated by a dead neighbor; no feature to take, one lesson to keep.

Integration candidates

C1 — ADR Confirmation section + YAML frontmatter (MADR 4.0 parity). Add an optional ## Confirmation section to adr-template.md and optional YAML frontmatter (status, date, deciders). Fence check: additive to existing .architecture/ dirs (Constraints ✓), deepens the decision→recalibration→progress-tracking spine (Purpose ✓), markdown-only (✓). Frontmatter also gives the MCP server's ADR listing something machine-readable to filter on without parsing headings. Smallest, most certain candidate — this is a docs-fix-class PR if there's appetite.

C2 — Scale-adaptive review depth (BMAD's one good idea, minus the agile machinery). Today architecture-review convenes the full board regardless of change size. Add review tiers: a quick single-specialist pass for small changes, full multi-perspective board for architectural ones, selected by explicit user choice (not magic heuristics). Fence check: deepens the review differentiator rather than widening surface (Direction ✓); no new dependencies (✓). Complements #35/#36, doesn't overlap them — those make personas better, this decides how many to summon.

C3 — Single-source USAGE generation (spec-kit's maintenance lesson). We hand-maintain five USAGE-.md variants; drift between them is a standing kill-criterion in this repo's own review panel. spec-kit sustains 30+ integrations only because per-assistant files are generated from one source. A small build step (or even a checked-in tools/ script + CI drift check) that renders the shared sections of USAGE-.md from one template would convert a recurring bug class into a mechanical impossibility. Fence check: "tracking the extension surface is maintenance" (Direction ✓); tooling stays optional — the generated markdown is what ships (Constraints ✓).

C4 — Principles as an explicit review gate (constitution parity). spec-kit's constitution works because every phase must reference it. We ship .architecture/principles.md but reviewers aren't required to check findings against it. Add a step to the review skill: each specialist flags any finding that conflicts with a recorded principle, and any recommendation that would change a principle gets called out as such. Fence check: deepens reviews (Direction ✓), zero new surface (✓).

Non-candidates (recorded so they stay settled)

  • OpenSpec's delta-spec model — replacing or supplementing append-only ADRs with living current-state specs is a different institutional-memory philosophy. Our Constraints make ADR history sacred and migrations additive; adopting deltas would fork the mental model. Wont-pursue unless user demand shows up.
  • log4brains-style knowledge-base app or site generator — runtime tooling around the markdown is exactly what decayed at log4brains. At most a future docs paragraph on publishing .architecture/ with any SSG; no code.
  • Chasing spec-kit's 30-assistant matrix — anti-goal verbatim: an integration must bring evidence of demand. Codex/Cursor stay best-effort.

Intended labels: loop:research, status:analyzed.

Sources: spec-kit · spec-kit review 2026 · BMAD-METHOD · BMAD v6 overview · OpenSpec · MADR · MADR template primer · log4brains

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    loop:researchCompetitive / related-project research loopstatus:analyzedAnalyzed, awaiting decision

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions