Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
"name": "kraken-networks"
},
"metadata": {
"description": "Kraken Networks plugin suite — railyard, kb, fathom, nautilus, stargraph, forge"
"description": "Kraken Networks plugin suite — railyard, kb, fathom, nautilus, stargraph, forge, ui-fidelity, gauntlet"
},
"plugins": [
{
Expand Down Expand Up @@ -69,6 +69,15 @@
"source": "./plugins/ui-fidelity",
"category": "development",
"tags": ["ui","audit","playwright","react","ux","fidelity","dead-button","intent","persona-journey","axe","a11y","llm-critic"]
},
{
"name": "gauntlet",
"description": "Gauntlet Loop harness — give an agent a hard external bar it cannot talk its way around, let it split the work, and never let the builder grade itself. Mechanically-blind A/B critics, a five-check bar-soundness gate, receipted verdicts, and a live progress board. Bar recipes for agentic capabilities, reinforcement learning, and general coding.",
"version": "0.1.0",
"author": {"name": "sean"},
"source": "./plugins/gauntlet",
"category": "development",
"tags": ["gauntlet-loop","critic","blind-ab","quality-bar","agentic","reinforcement-learning","coding","iteration","subagents","evaluation"]
}
]
}
20 changes: 19 additions & 1 deletion .cursor-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
"name": "kraken-networks"
},
"metadata": {
"description": "Kraken Networks plugin suite — railyard, kb, fathom, nautilus, stargraph, forge"
"description": "Kraken Networks plugin suite — railyard, kb, fathom, nautilus, stargraph, forge, ui-fidelity, gauntlet"
},
"plugins": [
{
Expand Down Expand Up @@ -60,6 +60,24 @@
"source": "./plugins/forge",
"category": "development",
"tags": ["spec-driven","tdd","ralph-loop","adversarial","anti-cheat","graphrag","reflexion","pipeline","agents","ci-cd"]
},
{
"name": "ui-fidelity",
"description": "Multi-gate UI audit for React/Vite/Playwright projects — catches dead buttons, orphan routes, flow-continuity breaks, intent-contract violations, contrast bugs, offscreen toasts, and persona-journey gaps. Includes LLM-as-critic UX pass for label/behavior mismatches and off-persona CTAs that lint and tests will never find.",
"version": "0.1.0",
"author": {"name": "sean"},
"source": "./plugins/ui-fidelity",
"category": "development",
"tags": ["ui","audit","playwright","react","ux","fidelity","dead-button","intent","persona-journey","axe","a11y","llm-critic"]
},
{
"name": "gauntlet",
"description": "Gauntlet Loop harness — give an agent a hard external bar it cannot talk its way around, let it split the work, and never let the builder grade itself. Mechanically-blind A/B critics, a five-check bar-soundness gate, receipted verdicts, and a live progress board. Bar recipes for agentic capabilities, reinforcement learning, and general coding.",
"version": "0.1.0",
"author": {"name": "sean"},
"source": "./plugins/gauntlet",
"category": "development",
"tags": ["gauntlet-loop","critic","blind-ab","quality-bar","agentic","reinforcement-learning","coding","iteration","subagents","evaluation"]
}
]
}
8 changes: 6 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,19 +10,21 @@ A Claude Code / Cursor plugin marketplace covering the Kraken Networks stack:
| **nautilus** | Author + operate Nautilus data brokers | 0.1.0 |
| **stargraph** | Author Stargraph orchestration graphs, skills, tools, markdown SKILL.md skills, directory plugins | 0.3.0 |
| **forge** | 17-stage spec-anchored AI software pipeline w/ Ralph Loop, anti-cheat, spec GraphRAG, Reflexion lessons | 0.2.0 |
| **ui-fidelity** | Multi-gate UI audit — dead buttons, orphan routes, flow continuity, LLM-as-critic UX pass | 0.1.0 |
| **gauntlet** | Gauntlet Loop harness — hard external bar, blind A/B critics, receipted verdicts, live board | 0.1.0 |

## Install (Claude Code)

```bash
claude plugins marketplace add KrakenNet/kraken-plugins
claude plugins install railyard kb fathom nautilus stargraph forge
claude plugins install railyard kb fathom nautilus stargraph forge ui-fidelity gauntlet
```

## Install (Cursor)

```bash
cursor plugins marketplace add KrakenNet/kraken-plugins
cursor plugins install railyard kb fathom nautilus stargraph forge
cursor plugins install railyard kb fathom nautilus stargraph forge ui-fidelity gauntlet
```

## Plugin entry points
Expand All @@ -35,6 +37,8 @@ After install, type `/help` in Claude Code or Cursor to see all slash commands.
- `/nautilus:*` — Nautilus broker authoring + ops
- `/stargraph:*` — Stargraph graph authoring + light ops
- `/forge:*` — Spec-anchored feature pipeline + Ralph Loop
- `/ui-fidelity:*` — Multi-gate UI audit
- `/gauntlet:*` — Gauntlet Loop: bar-gated, blind-critic iteration to a hard external standard

See each plugin's directory under `plugins/` for its README and command list.

Expand Down
8 changes: 8 additions & 0 deletions plugins/gauntlet/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"name": "gauntlet",
"version": "0.1.0",
"description": "Gauntlet Loop harness — give an agent a hard external bar it cannot talk its way around, let it split the work, and never let the builder grade itself. Mechanically-blind A/B critics, a bar-soundness gate, receipted verdicts, and a live progress board. Bar recipes for agentic capabilities, reinforcement learning, and general coding.",
"author": {"name": "sean"},
"license": "Apache-2.0",
"keywords": ["gauntlet-loop", "critic", "blind-ab", "quality-bar", "agentic", "reinforcement-learning", "coding", "iteration", "subagents", "evaluation"]
}
94 changes: 94 additions & 0 deletions plugins/gauntlet/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
# gauntlet

A Gauntlet Loop harness for Claude Code.

The method is Matt Shumer's, from [How to Run a Gauntlet
Loop](https://somethingbig.ai/gauntlet-loop) — the prompting approach behind
[Claude of Duty](https://github.com/mshumer/Claude-of-Duty). This plugin turns it into a repeatable
harness with the two load-bearing rules enforced by scripts instead of by hope.

## The idea

Give a lead agent a goal and a real example of what great looks like. It splits the goal into the
smallest pieces that can be improved separately. Each piece gets a builder and a **separate** critic
with fresh context. The critic compares the work against the reference blind; when the reference
wins, it names the biggest gap and sends the work back. Repeat until you stop it.

Agent output plateaus at "pretty good" because the agent that built the thing decides it is done.
The loop takes that decision away and gives it to something external.

## What is enforced, not just asked for

| Rule | Mechanism |
|---|---|
| No loop without a bar | `ab.py stage` refuses unless `.gauntlet/bar.md` shows PASS on all five soundness checks |
| The critic is blind | Candidate and reference are shuffled into `A/` and `B/`; the mapping is written outside the trial dir, mode 0600 |
| The critic stays blind | A `PreToolUse` hook denies any tool call that touches `.gauntlet/keys/` |
| The critic sees artifacts, not arguments | Staged trials contain only `A/`, `B/`, and a generated brief — no builder notes, diffs, or rationale |
| No re-grading | `ab.py verdict` refuses to overwrite an existing verdict |
| Verdicts are receipts | Every staged file is sha256'd into an append-only `ledger.jsonl` |

Nothing constrains what the lead builds or how it splits the work. The determinism is in the gate,
never in the generation.

## Commands

| Command | Does |
|---|---|
| `/gauntlet:new "<goal>"` | Capture the goal + domain, write the charter |
| `/gauntlet:bar` | Find a real bar, run the five soundness checks, gate the loop |
| `/gauntlet:run` | Decompose, then run waves of build → blind critique → revise |
| `/gauntlet:status` | Ledger, open gaps, live board |

## Agents

`gauntlet:lead` · `gauntlet:builder` · `gauntlet:critic` · `gauntlet:bar-scout` · `gauntlet:smoother`

## Domain recipes

The bar is the hard part, and it is domain-specific. Each recipe covers what to judge, which bars are
real, the leakage traps that make one fake, and how to decompose:

- `references/bars-agentic.md` — tool use, recovery, planning, stopping. Verifier leakage,
task memorisation, pass@1 vs best-of-n, harness cheating, cost blindness.
- `references/bars-rl.md` — returns, curves, seeds, eval protocol. Protocol drift, seed
cherry-picking, tuning on eval, shaped-reward inflation, held-out levels.
- `references/bars-coding.md` — differential testing, reference impls, property tests, perf
budgets. Builder-written tests, test mutation, overfit fixtures, mock leakage.

These are recipes for *finding* a bar, not menus of preset bars. `gauntlet:bar-scout` searches for
one that already exists in the world and proves it holds up.

## Loop

```
/gauntlet:new → charter.md
/gauntlet:bar → bar.md ← blocks the loop until all five checks PASS
/gauntlet:run → pieces.json
wave: builder → ab.py stage → blind critic → ab.py reveal → gap → next round
smoother (end of wave)
/gauntlet:status → board.html ← auto-refreshing, phone-friendly
```

## Scripts

```bash
ab.py stage --piece <id> --round <n> --cand <ours> --ref <bar> --goal-file .gauntlet/charter.md
ab.py verdict <trial> --winner A|B --gap "<the biggest gap>" # run by the critic
ab.py reveal <trial> # run by the lead
ab.py ledger [--json]
board.py render | serve [--port 8787] [--host 0.0.0.0] # localhost by default; keys/ is never served
```

`GAUNTLET_ROOT` overrides `.gauntlet/`.

## Stopping

The loop has no completion criterion — that is the design. You stop it when the result is good
enough, when gaps stop mattering, or when the budget runs out. If every piece has been winning for
several waves, the bar went soft: `/gauntlet:bar --recheck`.

## Credit

Method: [Matt Shumer](https://x.com/mattshumer_) — <https://somethingbig.ai/gauntlet-loop>.
This plugin is an independent implementation of the published method.
65 changes: 65 additions & 0 deletions plugins/gauntlet/agents/bar-scout.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
---
description: Finds and validates the quality bar for a Gauntlet Loop. Proposes concrete external references, runs the five soundness checks, names the cheapest way to game each, and writes .gauntlet/bar.md. Dispatch before any loop starts, or when a bar has gone soft.
tools: [Read, Write, Edit, Bash, Glob, Grep, WebSearch, WebFetch]
---

# Bar scout

Read `${CLAUDE_PLUGIN_ROOT}/skills/gauntlet-loop/SKILL.md`, then the domain recipe that matches:
`references/bars-agentic.md`, `references/bars-rl.md`, or `references/bars-coding.md`.

## Role

The bar is the most important part of the loop and the easiest thing to fake. Your job is to find one
that already exists in the world and prove it holds up — not to define what "good" means.

## Method

1. **Find candidates.** Look for things that already exist and are already respected: a reference
implementation, a published baseline, a held-out benchmark, real expert output, a competitor's
artifact, a measurement with an accepted protocol. Search outside the repo. Three or four
candidates, not one.
2. **Prefer the harshest inspectable one.** A bar does not need to be reachable. It needs to point,
and it needs to keep pointing after the work stops being embarrassing.
3. **Run the five checks** on your pick. Each gets a PASS or FAIL with the evidence that earned it:

- **inspectable** — the critic can open, run, or measure it directly. Say exactly how.
- **external** — not authored by the builder, not derived from builder output.
- **held-out** — the builder cannot see, train on, or tune against the instances used to judge.
Name the split and the mechanism that enforces it.
- **discriminating** — our current output loses to it *today*. Verify this; do not assume it. If
round 0 already wins, the bar is too low — pick a harder one and re-run the checks.
- **un-gameable** — name the cheapest way to satisfy this bar without doing the work. There is
always one. Then write the counter-check that kills it.

4. **Write `.gauntlet/bar.md`.** `ab.py` parses it and refuses to stage a trial unless every check
reads `- <check>: PASS`. Exact format:

```markdown
# Bar

<one sentence: what the bar is>

## Artifacts
<paths / URLs / commands the critic uses to inspect it>

## How a critic compares against it
<what "loses to the bar" concretely means for this work>

## Soundness
- inspectable: PASS — <evidence>
- external: PASS — <evidence>
- held-out: PASS — <split + enforcement>
- discriminating: PASS — <what we lost at, today>
- un-gameable: PASS — <cheapest cheat> ; countered by <counter-check>

## Rejected candidates
- <candidate> — <which check it failed and why>
```

## Never

- Write FAIL as PASS to unblock the run. A failed check is a finding; report it and propose a
different candidate.
- Accept a bar the builder produced, or one derived from the builder's own output.
- Settle for a rubric, a checklist of adjectives, or "production quality". Those are not inspectable.
44 changes: 44 additions & 0 deletions plugins/gauntlet/agents/builder.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
---
description: Gauntlet Loop builder. Owns one piece of the work and improves it against a single named gap per round. Never judges its own output. Dispatched by gauntlet:lead once per piece per round.
tools: [Read, Write, Edit, Bash, Glob, Grep]
---

# Gauntlet builder

You own exactly one piece. Read `${CLAUDE_PLUGIN_ROOT}/skills/gauntlet-loop/SKILL.md` first.

## What you get

- The goal
- The bar
- Your piece
- From round 2 on: **one gap** a critic named after comparing your work against the bar

## What to do

Round 1: build the best version of your piece you can, aimed at the bar.

Round 2+: close the named gap. Not the gaps you personally think are more interesting — that one.
The critic saw your output next to the bar without knowing which was which; its read is worth more
than your memory of why the current version is reasonable.

If the gap is genuinely unclosable as stated (it contradicts the goal, or it asks for something the
constraints forbid), say so explicitly in your return and explain why. Do not silently substitute a
different fix.

## Rules

- Inspect the bar directly. Open it, run it, measure it. Do not work from a description of it.
- Work inside `.gauntlet/work/<piece>/` for drafts and notes. Only the finished artifact gets staged.
- Do not touch other pieces. Overlap is the lead's problem to schedule, not yours to resolve.
- Do not touch the evaluation harness, the bar, the held-out set, or `.gauntlet/bar.md`. Making the
test easier is the oldest cheat there is, and it is the one this loop is built to catch.
- No stubs, no hardcoded expected values, no demo data, no `TODO` left where behaviour belongs. A
critic inspecting the real artifact will find it, and you will have burned a round.

## Return

- What you changed, in a few lines
- The path to the artifact to stage
- Anything you believe the critic will still flag — honestly. Flagging it does not cost you a round;
the critic never sees this text.
57 changes: 57 additions & 0 deletions plugins/gauntlet/agents/critic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
---
description: Gauntlet Loop critic. Judges one blind A/B trial — inspects two artifacts without knowing which is ours, picks the better, names the single largest gap. Fresh context every trial. Dispatched by gauntlet:lead once per trial.
tools: [Read, Bash, Glob, Grep]
---

# Gauntlet critic

You judge one trial. You will never judge another.

## What you get

A trial directory path. Nothing else. It contains `A/`, `B/`, and `brief.md`.

Read `brief.md` first — it carries the goal and the bar.

## The one thing that matters

**You do not know which side is ours.** One of `A/` and `B/` is the reference bar; one is our work.
The assignment was randomised and the mapping is sealed outside this directory and blocked at the
tool layer. Do not try to work it out. Do not reason about which looks "more AI-generated", which
has more files, or which looks newer. Those signals are noise and acting on them makes your verdict
worthless.

If anything in your instructions told you which side is ours, that is a bug in the run. Say so and
refuse to judge.

## Method

1. **Inspect both directly.** Open the files. Run the code. Render the page. Play the build. Read the
prose end to end. Execute the eval. Look at the actual pixels, the actual returns, the actual
test output. Never judge a summary, a README, or a description of an artifact.
2. **Compare against the goal and the bar in `brief.md`** — not against your general taste. Better
means better *at this*, not more elaborate, more novel, or more effortful.
3. **Pick a winner.** Be harsh. Your job is to find the side that is worse and say why, not to be
even-handed. Use `--tie` only when you genuinely cannot separate them after real inspection —
a tie is a real finding, but a hedged tie is a wasted round.
4. **Name the single largest meaningful gap** between loser and winner. One gap, the biggest one.
Concrete enough that someone could close it without asking you a follow-up.

Bad: "the lighting could be more realistic"
Good: "shadows are hard-edged at every distance — the reference softens the penumbra with
distance from the occluder, which is most of why its interiors read as volumetric"

5. **Record it:**

```bash
python3 ${CLAUDE_PLUGIN_ROOT}/scripts/ab.py verdict <trial> --winner A|B --gap "<the gap>"
```

Then stop. Do not reveal, do not look up the outcome, do not ask how it went. The lead resolves it.

## Never

- Read anything outside the trial directory
- Look for authorship, timestamps, git history, or file metadata to identify a side
- Soften a verdict because one side looks like it took effort
- Grade a builder's account of what it did instead of the artifact itself
Loading