Repository navigation
feat(evals): add build-cli-004-worktree-stacks (CLI-2400) - #307
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
f4af609 to
9e35200
Compare
c883a62 to
3158ffd
Compare
|
Heads-up: I rebased this branch ( Why: #281 was split into two PRs this morning — #308 (the Nothing you need to do. Your eval is already in Two things from the review round on #308 that touch your runs, for awareness only:
|
Final chunk of the closed #308 — the piece @mattrossman asked for: the minimal CLI suite plumbing and CODEOWNERS, standalone, with no CLI experiment internals to review. **15 files, 260 lines**, and four of those files are 11–13 lines each. Everything it depended on is now on `main` (#315's `experiments/<owner>/` layout, and #324 + #325's sandbox options), so this is based on `main` and the diff is only its own content. ## What it adds `cli` as an eval suite and an experiment suite, `@supabase/cli` ownership over `/evals/cli/`, `/experiments/cli/` and `cli-eval-results.json`, and the workflow wiring so `eval-refresh` and `append-gh-pages-history` know the suite exists. The substance is five environment columns: one scenario run unchanged across forced CLI environments, so a failure isolates to *which* environment broke. They're thin now that `experiments/presets.ts` exists — the whole of the Docker-less column is: ```ts export default defineExperiment({ ...codexGpt6Luna, suite: ['cli'], // beta: the Docker-less path only exists in the managed stack, which ships in beta. localStack: localStackRuntime({ cliVersion: 'beta', docker: 'absent' }), skipEval: skipUnlessDockerless, }); ``` | column | environment | picks up | |---|---|---| | `codex-gpt-6-luna-cli-pinned` | repo-pinned CLI, Docker available | every `interface: cli` eval that isn't hosted-linked | | `…-cli-stable` | npm `latest` | same | | `…-cli-beta` | npm `beta` | same | | `…-cli-nodaemon` | beta, daemon unreachable | also needs `needsDocker: false` + `projectRunning: false` | | `…-cli-absent` | beta, no `docker` binary | also needs `needsDocker: false` + `projectRunning: false` | All five columns share one `skipEval: skipUnlessCli`, so they run the same eval set and stay comparable — which is the whole point of the suite. An earlier revision made the pinned column the shared benchmark experiment with `'cli'` appended to its `suite`; that experiment has no `skipEval`, so it picked up hosted-linked evals the other four skip. A CLI-owned pinned experiment fixes that and leaves this PR touching nothing outside CLI-owned paths. `skipUnlessCli` / `skipUnlessDockerless` are the only things left over from the old `experiments/_lib/`; they live in `experiments/cli/lib/`, which discovery ignores for free since it only matches `*.experiment.ts` directly under an owner directory. ## Verified, including the parts tests can't reach `format:check`, `typecheck`, and the sandbox, core, framework, `test:cli-lib`, vercel-runner and web suites all pass. Two things unit tests can't cover, checked directly instead: - **Discovery**, since it's filename-convention-based and fails silently: all 16 experiments resolve and load, `experiments/cli/lib/` is correctly not treated as an experiment, and the five columns resolve to the intended runtimes and channels (`beta`, `beta`, `beta`, `stable`, and none for the pinned baseline). - **The workflow**, which only ever runs in CI, no longer resolves channels at all — it passes through the manual `cli_stable_version` / `cli_beta_version` dispatch inputs and otherwise lets `run-vercel-evals.ts` derive the channels from the actual pairs. Pre-filling both env vars would have short-circuited that narrowing. Each run's results now record the CLI version the sandbox actually installed, read from the session's environment marker, so a `stable` row names its binary rather than just its channel. That needed `cliVersionSchema` widened — it rejected every prerelease, so a beta run's real version could not have been recorded at all. `test:cli-lib` carries `--passWithNoTests` deliberately: `evals/cli/` has no scorer tests until #281/#314/#316 land, and without the flag the script's exit code depends on which of its two paths happens to be populated. ## Note on the column names These follow the `gpt-6-luna` rename from #328. They key `cli-eval-results.json`, and no results exist yet, so there is nothing to churn — but that also means the names should settle before the suite is first run. ## Next #281, #314, #316 and #307 get re-parented onto `main` once this lands — they currently point at the retained branch of the closed #308. /cc @Rodriguespn @kanadgupta
…2400) One human plus N coding agents, each in its own git worktree, is the headline workflow for the Slim CLI launch (FDBKIN-20391). Each worktree needs its own isolated local Supabase stack with zero leakage between them. Nothing in this repo measured that. Scenario: starting from an empty sandbox, the agent creates a repo with worktrees feature-a/b/c, starts a local stack in each, and adds a different table (widgets/gadgets/gizmos) plus one seed row per worktree. No seed data and no framework changes: the agent builds everything, including the worktrees. Scorer (end state only, all via ctx.exec because the harness's built-in ctx.query/stackStatus resolve a single stack from the workspace root): - three real git worktrees on distinct branches, discovered by finding a repo in the workspace and asking `git worktree list --porcelain`, so worktrees the agent placed outside the workspace still count and three plain directories don't; - one live stack per worktree with three distinct host:port database endpoints, which is what catches aliasing/reuse; - per-table schema isolation via to_regclass against all three stacks; - at least one row per table in its home stack; - each table created by a migration file in its worktree; - an always-passing metrics check reporting fleet wall-clock across the three stacks, CLI version/channel, per-stack backend and runtime, and the number of `supabase start` invocations. Stack resolution asks the managed backend first (SUPABASE_EXPERIMENTAL_STACK=1 supabase stack status --env) and falls back to the legacy one, the same shape as build-database-002-stack-lifecycle. Verified against CLI 2.118.0-beta.37: the managed backend rejects the legacy -o flag, the DB password is random per stack, and the legacy backend cannot see managed stacks. Suite placement follows #281: `evals/cli/` with `needsDocker: false`, so the eval runs on the pinned Codex Luna baseline plus the four `codex-gpt-5.6-luna-cli-*` arms with no experiment edits. `services:` is omitted because the sandbox shim's legacy `-x` names are rejected by the managed backend. Two details come from the first CI run of an earlier draft (#295, 18 runs across six experiments): - The prompt now says the project is brand-new and asks for the worktrees as folders in this directory. The earlier "set this repo up" made 12 of 18 agents look for a repository in the empty workspace and stop to ask for one within about ten seconds. - Worktree discovery goes through git rather than a name search under the workspace. Two agents built everything correctly with worktrees at /tmp/feature-*, and the name search reported them as missing. The two runs that passed did so on the legacy backend by hand-editing project_id and every port per worktree; the metrics check records the backend so that path stays distinguishable from a managed-stack pass. Pure helpers are unit-tested in scoring.test.ts (run hint at the top of the file); the shapes in the fixtures were captured from real beta CLI output. The full scorer was also run against three hand-built native stacks on the beta CLI, with worktrees both inside and outside the workspace, and passes all checks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…entMarker()
Replace the local /tmp/supabase-eval-runtime.json reader with the
scoring context's environmentMarker(), which reads as root and
validates the marker's shape. Update the README's stale
codex-gpt-5.6-luna-cli-{stable,beta,nodaemon,absent} arm names to the
current five codex-gpt-6-luna-cli-{pinned,stable,beta,nodaemon,absent}
environments, and correct the claim that pinned is a separate shared
baseline rather than one of the five. Drop leftover references to the
removed experiments/_lib/docker-aware-local-stack.ts helper.
bdc011c to
c3c6459
Compare
1974f30 to
365772b
Compare
# Conflicts: # apps/web/src/data/cli-eval-results.json
Coly010
left a comment
There was a problem hiding this comment.
Hey kanad 🙂 I merged main into this to clear the results conflict. It's a merge commit, no force-push, so nothing should've moved under you. I kept main's results and appended your 15 build-cli-004 rows; they'll get replaced once the refresh that's running now lands
I did a fairly adversarial pass on the scorer, trying both to fake a pass and to get a valid run failed, and a few things came out of it. Sorry it's a long one!
The bigger ones:
scoring.tsduplicates a fair bit ofevals/cli/lib, and a couple of the copies are weaker than the lib versions. That's where most of the false fails below come from. I had the same thing on resolve-database-003, and rebuilding it on the lib (3286d9b) shrank it a lot, so hopefully it's less painful than it looks- "tell me how to reach each stack" isn't scored
- the migration check can pass without a migration actually being applied
- the committed results predate #355, so nodaemon's 0/3 is effectively the same as absent. Worth waiting for the refresh before reading much into the Docker-less arms. Also interesting that no run passed via a managed stack; every pass was legacy with hand-edited ports
The PR body numbers are a bit stale too: 5×2 is now 5×5, 28 cases is 29, the "Stacked on #281" note no longer applies, and the #295 failures add up to 14 of 16. Probably easiest to swap the #295 history for the post-refresh numbers once they're in
Details are inline. The nits are genuinely take-it-or-leave-it, and if you disagree with any of it that's totally fine. I've got all the counterexamples as vitest cases if you want them as regression tests. Otherwise this is looking really close!
# Conflicts: # apps/web/src/data/cli-eval-results.json
Split the monolithic scoring.ts into domain modules with their own tests and keep composition, the detour judge and the default export in EVAL.ts. Resolve each worktree's stack with resolveStackWithAgentHomes (named stacks and relocated homes), read migrations through a shared evals/cli/lib/migrations.ts that build-database-002 now uses too, score the reported database ports, count starts from parsed invocations, pick the repo deterministically, and align the prompt and README with the scorer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tracks CLI-2400 under the Slim CLI evals RFC. Supersedes #295.
The problem it's solving
The headline agent workflow for the Slim CLI launch is one human plus N coding agents, each working in its own git worktree, each needing its own isolated local Supabase stack. Nothing in this repo measured whether an agent can get there: three worktrees, three stacks, zero leakage.
What the PR adds
One new eval in the
clisuite,evals/cli/build-cli-004-worktree-stacks/. Starting from an empty sandbox, the agent is asked to set up a repo with worktreesfeature-a,feature-b,feature-c, start a local stack in each, and add a different table (widgets/gadgets/gizmos) via a migration plus one sample row per worktree, then say how to reach each stack.needsDocker: falseputs it on the Docker-less arms, so it runs under all sixcodex-gpt-6-luna-cli-*experiments (pinned, stable, beta, next, nodaemon, absent).How it differs from its siblings:
build-database-003-parallel-projectsruns separate projects, each with its ownconfig.toml, andbuild-database-004-named-stacksruns named stacks from one project directory. Here the worktrees share one committedconfig.toml, so they collide on ports unless the agent finds the managed stack's automatic ports or edits the config per worktree.A scorer that grades the end state, never the method (17 checks).
EVAL.tsholds the composition and the judge call. The helpers live inworktrees.ts,stacks.ts,tables.ts,migrations.ts,report.tsandmetrics.ts, each with its own tests.git worktree list --porcelainand the best match wins, so worktrees placed outside the workspace still count and a stray repo can't shadow the real one;host:portdatabase endpoints. Stacks resolve throughresolveStackWithAgentHomes, sosupabase stack start --stack feature-aand a per-worktreeSUPABASE_HOMEare both found;to_regclassagainst all three stacks, and at least one row per table in its home stack;supabase_migrations.schema_migrations;no container-runtime detoursjudge, the same asbuild-database-002/003/004, so installing Docker on a Docker-less arm and using the legacy start fails;metricscheck: fleet wall-clock across the three stacks, CLI version and channel, per-stack backend, runtime and relocated home,stack startvs legacystartcounts, and whetherSUPABASE_EXPERIMENTAL_STACKwas set.maskSqlLiteralsAndComments,checkMigrationAppliedand the table-creation matcher now live inevals/cli/lib/migrations.ts.build-database-002uses them from there with no behaviour change.Deliberate choices
services:is omitted, because the sandbox shim's legacy-x gotrue,kong,...is rejected by the managed backend.supabase initpins ports inconfig.toml. The managed stack treats them as exact intents, so the second worktree's start fails withPersisted database port is unavailableuntil the ports are removed. That is the "no port surgery" promise the launch makes.CI results (refresh
7951cde, 3 runs per experiment)12 of 18 pass, 9 of them on the managed stack with no port edits. Unlike #295, the Docker-less arms now pass. The six failures fall into two groups:
Verification
EVAL.test.ts. The counterexamples from review (commented or string-literalcreate table, never-applied migration,[task]noise before the JSON, named stacks, per-worktree homes, echoedsupabase start) are regression cases.pnpm format:check,pnpm typecheck,test:cli-lib, core/sandbox/web testspnpm eval:dry -- --suite cli --experiment-suite cliselects the eval on all six cli experimentsrun-evals-changedlabel (table above)🤖 Generated with Claude Code