Skip to content

feat(evals): add build-cli-004-worktree-stacks (CLI-2400) - #307

Merged
Coly010 merged 9 commits into
mainfrom
kanad-claude/cli-2400-worktree-stacks-cli-suite
Oct 9, 2026
Merged

Coly010 merged 9 commits into
mainfrom
kanad-claude/cli-2400-worktree-stacks-cli-suite

Conversation

@kanadgupta

@kanadgupta kanadgupta commented Sep 17, 2026 •

Copy link
Copy Markdown
Member

Tracks CLI-2400 under the Slim CLI evals RFC. Supersedes #295.

The problem it's solving

The headline agent workflow for the Slim CLI launch is one human plus N coding agents, each working in its own git worktree, each needing its own isolated local Supabase stack. Nothing in this repo measured whether an agent can get there: three worktrees, three stacks, zero leakage.

What the PR adds

One new eval in the cli suite, evals/cli/build-cli-004-worktree-stacks/. Starting from an empty sandbox, the agent is asked to set up a repo with worktrees feature-a, feature-b, feature-c, start a local stack in each, and add a different table (widgets / gadgets / gizmos) via a migration plus one sample row per worktree, then say how to reach each stack. needsDocker: false puts it on the Docker-less arms, so it runs under all six codex-gpt-6-luna-cli-* experiments (pinned, stable, beta, next, nodaemon, absent).

How it differs from its siblings: build-database-003-parallel-projects runs separate projects, each with its own config.toml, and build-database-004-named-stacks runs named stacks from one project directory. Here the worktrees share one committed config.toml, so they collide on ports unless the agent finds the managed stack's automatic ports or edits the config per worktree.

A scorer that grades the end state, never the method (17 checks). EVAL.ts holds the composition and the judge call. The helpers live in worktrees.ts, stacks.ts, tables.ts, migrations.ts, report.ts and metrics.ts, each with its own tests.

  • three real git worktrees on distinct branches. Every repo in the workspace is asked for git worktree list --porcelain and the best match wins, so worktrees placed outside the workspace still count and a stray repo can't shadow the real one;
  • one live stack per worktree with three distinct host:port database endpoints. Stacks resolve through resolveStackWithAgentHomes, so supabase stack start --stack feature-a and a per-worktree SUPABASE_HOME are both found;
  • per-table schema isolation via to_regclass against all three stacks, and at least one row per table in its home stack;
  • each table created by a migration file in its worktree, with comments and string literals masked, and that migration's version recorded in the worktree stack's supabase_migrations.schema_migrations;
  • the final message reports each worktree's real DB or API port;
  • the shared no container-runtime detours judge, the same as build-database-002/003/004, so installing Docker on a Docker-less arm and using the legacy start fails;
  • an always-passing metrics check: fleet wall-clock across the three stacks, CLI version and channel, per-stack backend, runtime and relocated home, stack start vs legacy start counts, and whether SUPABASE_EXPERIMENTAL_STACK was set.

maskSqlLiteralsAndComments, checkMigrationApplied and the table-creation matcher now live in evals/cli/lib/migrations.ts. build-database-002 uses them from there with no behaviour change.

Deliberate choices

  • Nothing tells the agent about the managed stack or its feature flag. Whether agents can discover it is part of what's measured.
  • services: is omitted, because the sandbox shim's legacy -x gotrue,kong,... is rejected by the managed backend.
  • Known footgun left in. supabase init pins ports in config.toml. The managed stack treats them as exact intents, so the second worktree's start fails with Persisted database port is unavailable until the ports are removed. That is the "no port surgery" promise the launch makes.

CI results (refresh 7951cde, 3 runs per experiment)

experiment CLI passed
pinned 2.117.0 2/3 (both legacy)
stable 2.120.0 1/3 (managed)
beta 2.121.0-beta.12 2/3 (1 managed, 1 legacy)
next 3.0.0-next.3 3/3 (managed)
nodaemon 2.121.0-beta.12 2/3 (managed)
absent 2.121.0-beta.12 2/3 (managed)

12 of 18 pass, 9 of them on the managed stack with no port edits. Unlike #295, the Docker-less arms now pass. The six failures fall into two groups:

  • 4 runs ended early (about 1.5 to 2.5 minutes, 10 to 17 tool calls) with no worktrees: three with a repo in the workspace and one with no repo at all.
  • 2 long runs (pinned r3, stable r2) did the work but had no stack running at scoring time. Worktrees, branches and migration files all pass, but every backend reports no stack: the legacy containers are gone, and on stable the managed owner is unavailable.

Verification

  • Scorer unit tests for every module, plus an end-to-end EVAL.test.ts. The counterexamples from review (commented or string-literal create table, never-applied migration, [task] noise before the JSON, named stacks, per-worktree homes, echoed supabase start) are regression cases.
  • pnpm format:check, pnpm typecheck, test:cli-lib, core/sandbox/web tests
  • pnpm eval:dry -- --suite cli --experiment-suite cli selects the eval on all six cli experiments
  • Agent-backed runs via the run-evals-changed label (table above)

🤖 Generated with Claude Code

@kanadgupta
kanadgupta requested a review from a team as a code owner September 17, 2026 20:31
@kanadgupta kanadgupta added the run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes label Sep 17, 2026
@vercel

vercel Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
evals Ready Ready Preview Oct 9, 2026 9:47am UTC

Request Review

@kanadgupta
kanadgupta changed the base branch from main to columferry/cli-2398-add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle September 17, 2026 21:02
@kanadgupta kanadgupta added run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes and removed run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes labels Sep 17, 2026
@Coly010
Coly010 force-pushed the columferry/cli-2398-add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle branch from f4af609 to 9e35200 Compare September 18, 2026 08:24
@Coly010
Coly010 force-pushed the kanad-claude/cli-2400-worktree-stacks-cli-suite branch from c883a62 to 3158ffd Compare September 18, 2026 08:35
@Coly010

Coly010 commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Heads-up: I rebased this branch (kanad-claude/cli-2400-worktree-stacks-cli-suite) with --force-with-lease and it now carries only your two commits (f3d48c7 eval, 3158ffd CI results) on top of #281's current head.

Why: #281 was split into two PRs this morning — #308 (the cli suite, needsDocker, experiments/_lib runtime and the -cli-* experiments) and #281 (the stack-lifecycle eval only, stacked on #308) — which rewrote the history this branch was built on, so the PR had started showing my superseded commits in its diff. Nothing in your commits changed; pnpm check, the scorer tests for both CLI evals (90) and pnpm eval:dry -- --suite cli --experiment-suite cli (5 experiments × 2 evals) all pass on the rebased branch.

Nothing you need to do. Your eval is already in evals/cli/ with needsDocker: false, which is exactly how the Docker-less arms now select evals (skipUnlessDockerless in experiments/_lib), so it runs under all five arms without the experiment allow-list edits from #295. Merge order is #308 → #281 → this.

Two things from the review round on #308 that touch your runs, for awareness only:

  • The first CI run of the cli suite (35274398896, the one that produced your cli-eval-results.json) lost every beta-channel pair because npm's beta tag pointed at 2.118.0-beta.52 while its GitHub release was still a draft (.deb 404). The resolver now verifies the asset and falls back to the newest published beta, and the workflow pins one version per run — so the next refresh should fill in the -cli-beta/-cli-nodaemon/-cli-absent rows.
  • The Docker-less arms no longer mount the Docker socket at all (mountDockerSocket: false) in addition to the DOCKER_HOST + shim layering.

Coly010 added a commit that referenced this pull request Sep 23, 2026
Final chunk of the closed #308 — the piece @mattrossman asked for: the
minimal CLI suite plumbing and CODEOWNERS, standalone, with no CLI
experiment internals to review. **15 files, 260 lines**, and four of
those files are 11–13 lines each.

Everything it depended on is now on `main` (#315's
`experiments/<owner>/` layout, and #324 + #325's sandbox options), so
this is based on `main` and the diff is only its own content.

## What it adds

`cli` as an eval suite and an experiment suite, `@supabase/cli`
ownership over `/evals/cli/`, `/experiments/cli/` and
`cli-eval-results.json`, and the workflow wiring so `eval-refresh` and
`append-gh-pages-history` know the suite exists.

The substance is five environment columns: one scenario run unchanged
across forced CLI environments, so a failure isolates to *which*
environment broke. They're thin now that `experiments/presets.ts` exists
— the whole of the Docker-less column is:

```ts
export default defineExperiment({
  ...codexGpt6Luna,
  suite: ['cli'],
  // beta: the Docker-less path only exists in the managed stack, which ships in beta.
  localStack: localStackRuntime({ cliVersion: 'beta', docker: 'absent' }),
  skipEval: skipUnlessDockerless,
});
```

| column | environment | picks up |
|---|---|---|
| `codex-gpt-6-luna-cli-pinned` | repo-pinned CLI, Docker available |
every `interface: cli` eval that isn't hosted-linked |
| `…-cli-stable` | npm `latest` | same |
| `…-cli-beta` | npm `beta` | same |
| `…-cli-nodaemon` | beta, daemon unreachable | also needs `needsDocker:
false` + `projectRunning: false` |
| `…-cli-absent` | beta, no `docker` binary | also needs `needsDocker:
false` + `projectRunning: false` |

All five columns share one `skipEval: skipUnlessCli`, so they run the
same eval set and stay comparable — which is the whole point of the
suite. An earlier revision made the pinned column the shared benchmark
experiment with `'cli'` appended to its `suite`; that experiment has no
`skipEval`, so it picked up hosted-linked evals the other four skip. A
CLI-owned pinned experiment fixes that and leaves this PR touching
nothing outside CLI-owned paths.

`skipUnlessCli` / `skipUnlessDockerless` are the only things left over
from the old `experiments/_lib/`; they live in `experiments/cli/lib/`,
which discovery ignores for free since it only matches `*.experiment.ts`
directly under an owner directory.

## Verified, including the parts tests can't reach

`format:check`, `typecheck`, and the sandbox, core, framework,
`test:cli-lib`, vercel-runner and web suites all pass.

Two things unit tests can't cover, checked directly instead:

- **Discovery**, since it's filename-convention-based and fails
silently: all 16 experiments resolve and load, `experiments/cli/lib/` is
correctly not treated as an experiment, and the five columns resolve to
the intended runtimes and channels (`beta`, `beta`, `beta`, `stable`,
and none for the pinned baseline).
- **The workflow**, which only ever runs in CI, no longer resolves
channels at all — it passes through the manual `cli_stable_version` /
`cli_beta_version` dispatch inputs and otherwise lets
`run-vercel-evals.ts` derive the channels from the actual pairs.
Pre-filling both env vars would have short-circuited that narrowing.

Each run's results now record the CLI version the sandbox actually
installed, read from the session's environment marker, so a `stable` row
names its binary rather than just its channel. That needed
`cliVersionSchema` widened — it rejected every prerelease, so a beta
run's real version could not have been recorded at all.

`test:cli-lib` carries `--passWithNoTests` deliberately: `evals/cli/`
has no scorer tests until #281/#314/#316 land, and without the flag the
script's exit code depends on which of its two paths happens to be
populated.

## Note on the column names

These follow the `gpt-6-luna` rename from #328. They key
`cli-eval-results.json`, and no results exist yet, so there is nothing
to churn — but that also means the names should settle before the suite
is first run.

## Next

#281, #314, #316 and #307 get re-parented onto `main` once this lands —
they currently point at the retained branch of the closed #308. /cc
@Rodriguespn @kanadgupta
kanadgupta and others added 2 commits September 23, 2026 15:33
…2400)

One human plus N coding agents, each in its own git worktree, is the
headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree
needs its own isolated local Supabase stack with zero leakage between
them. Nothing in this repo measured that.

Scenario: starting from an empty sandbox, the agent creates a repo with
worktrees feature-a/b/c, starts a local stack in each, and adds a
different table (widgets/gadgets/gizmos) plus one seed row per worktree.
No seed data and no framework changes: the agent builds everything,
including the worktrees.

Scorer (end state only, all via ctx.exec because the harness's built-in
ctx.query/stackStatus resolve a single stack from the workspace root):

- three real git worktrees on distinct branches, discovered by finding a
  repo in the workspace and asking `git worktree list --porcelain`, so
  worktrees the agent placed outside the workspace still count and three
  plain directories don't;
- one live stack per worktree with three distinct host:port database
  endpoints, which is what catches aliasing/reuse;
- per-table schema isolation via to_regclass against all three stacks;
- at least one row per table in its home stack;
- each table created by a migration file in its worktree;
- an always-passing metrics check reporting fleet wall-clock across the
  three stacks, CLI version/channel, per-stack backend and runtime, and
  the number of `supabase start` invocations.

Stack resolution asks the managed backend first
(SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back
to the legacy one, the same shape as build-database-002-stack-lifecycle.
Verified against CLI 2.118.0-beta.37: the managed backend rejects the
legacy -o flag, the DB password is random per stack, and the legacy
backend cannot see managed stacks.

Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so
the eval runs on the pinned Codex Luna baseline plus the four
`codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is
omitted because the sandbox shim's legacy `-x` names are rejected by the
managed backend.

Two details come from the first CI run of an earlier draft (#295, 18
runs across six experiments):

- The prompt now says the project is brand-new and asks for the worktrees
  as folders in this directory. The earlier "set this repo up" made 12 of
  18 agents look for a repository in the empty workspace and stop to ask
  for one within about ten seconds.
- Worktree discovery goes through git rather than a name search under the
  workspace. Two agents built everything correctly with worktrees at
  /tmp/feature-*, and the name search reported them as missing.

The two runs that passed did so on the legacy backend by hand-editing
project_id and every port per worktree; the metrics check records the
backend so that path stays distinguishable from a managed-stack pass.

Pure helpers are unit-tested in scoring.test.ts (run hint at the top of
the file); the shapes in the fixtures were captured from real beta CLI
output. The full scorer was also run against three hand-built native
stacks on the beta CLI, with worktrees both inside and outside the
workspace, and passes all checks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…entMarker()

Replace the local /tmp/supabase-eval-runtime.json reader with the
scoring context's environmentMarker(), which reads as root and
validates the marker's shape. Update the README's stale
codex-gpt-5.6-luna-cli-{stable,beta,nodaemon,absent} arm names to the
current five codex-gpt-6-luna-cli-{pinned,stable,beta,nodaemon,absent}
environments, and correct the claim that pinned is a separate shared
baseline rather than one of the five. Drop leftover references to the
removed experiments/_lib/docker-aware-local-stack.ts helper.
@Coly010
Coly010 force-pushed the columferry/cli-2398-add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle branch from bdc011c to c3c6459 Compare September 23, 2026 15:05
@Coly010
Coly010 force-pushed the kanad-claude/cli-2400-worktree-stacks-cli-suite branch from 1974f30 to 365772b Compare September 23, 2026 15:05
@Coly010
Coly010 changed the base branch from columferry/cli-2398-add-a-docker-less-local-stack-e2e-eval-agent-cli-lifecycle to main September 23, 2026 15:05
# Conflicts:
#	apps/web/src/data/cli-eval-results.json

@Coly010 Coly010 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey kanad 🙂 I merged main into this to clear the results conflict. It's a merge commit, no force-push, so nothing should've moved under you. I kept main's results and appended your 15 build-cli-004 rows; they'll get replaced once the refresh that's running now lands

I did a fairly adversarial pass on the scorer, trying both to fake a pass and to get a valid run failed, and a few things came out of it. Sorry it's a long one!

The bigger ones:

  1. scoring.ts duplicates a fair bit of evals/cli/lib, and a couple of the copies are weaker than the lib versions. That's where most of the false fails below come from. I had the same thing on resolve-database-003, and rebuilding it on the lib (3286d9b) shrank it a lot, so hopefully it's less painful than it looks
  2. "tell me how to reach each stack" isn't scored
  3. the migration check can pass without a migration actually being applied
  4. the committed results predate #355, so nodaemon's 0/3 is effectively the same as absent. Worth waiting for the refresh before reading much into the Docker-less arms. Also interesting that no run passed via a managed stack; every pass was legacy with hand-edited ports

The PR body numbers are a bit stale too: 5×2 is now 5×5, 28 cases is 29, the "Stacked on #281" note no longer applies, and the #295 failures add up to 14 of 16. Probably easiest to swap the #295 history for the post-refresh numbers once they're in

Details are inline. The nits are genuinely take-it-or-leave-it, and if you disagree with any of it that's totally fine. I've got all the counterexamples as vitest cases if you want them as regression tests. Otherwise this is looking really close!

Comment thread evals/cli/build-cli-004-worktree-stacks/EVAL.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/PROMPT.md Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/PROMPT.md Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/scoring.ts Outdated
Comment thread evals/cli/build-cli-004-worktree-stacks/README.md
Coly010 and others added 3 commits October 9, 2026 10:22
# Conflicts:
#	apps/web/src/data/cli-eval-results.json
Split the monolithic scoring.ts into domain modules with their own tests and
keep composition, the detour judge and the default export in EVAL.ts. Resolve
each worktree's stack with resolveStackWithAgentHomes (named stacks and relocated
homes), read migrations through a shared evals/cli/lib/migrations.ts that
build-database-002 now uses too, score the reported database ports, count starts
from parsed invocations, pick the repo deterministically, and align the prompt
and README with the scorer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Coly010
Coly010 merged commit 8e7b1b3 into main Oct 9, 2026
10 checks passed

This branch was successfully deployed

1 active deployment
Preview — 7951cde6 Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-evals-changed Add to a PR to refresh only the benchmark evals that have had changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants