Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 0 additions & 33 deletions .github/workflows/eval-refresh.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,18 +41,6 @@ on:
type: boolean
required: false
default: false
cli_stable_version:
description: "Pin the cli suite's stable CLI version instead of resolving npm's latest dist-tag (blank to resolve)"
required: false
default: ""
cli_beta_version:
description: "Pin the cli suite's beta CLI version instead of resolving npm's beta dist-tag (blank to resolve)"
required: false
default: ""
cli_next_version:
description: "Pin the cli suite's next CLI version instead of resolving npm's next dist-tag (blank to resolve)"
required: false
default: ""
schedule:
- cron: '15 6 * * *'
pull_request:
Expand Down Expand Up @@ -94,9 +82,6 @@ jobs:
sandbox_concurrency: ${{ steps.inputs.outputs.sandbox_concurrency }}
filter_changed: ${{ steps.inputs.outputs.filter_changed }}
do_merge: ${{ steps.inputs.outputs.do_merge }}
cli_stable_version: ${{ steps.inputs.outputs.cli_stable_version }}
cli_beta_version: ${{ steps.inputs.outputs.cli_beta_version }}
cli_next_version: ${{ steps.inputs.outputs.cli_next_version }}
steps:
- name: Prepare inputs
id: inputs
Expand All @@ -112,9 +97,6 @@ jobs:
runs="${{ inputs.runs }}"
timeout_sec="${{ inputs.timeout_sec }}"
sandbox_concurrency="${{ inputs.sandbox_concurrency }}"
cli_stable_version="${{ inputs.cli_stable_version }}"
cli_beta_version="${{ inputs.cli_beta_version }}"
cli_next_version="${{ inputs.cli_next_version }}"

if [ -z "$suite" ] && [ -z "$eval_id" ]; then
echo "::error::Set suite or eval to choose which evals to run."
Expand All @@ -132,9 +114,6 @@ jobs:
runs="3"
timeout_sec="720"
sandbox_concurrency="250"
cli_stable_version=""
cli_beta_version=""
cli_next_version=""
else
experiments_override=""
eval_id=""
Expand All @@ -143,9 +122,6 @@ jobs:
runs="3"
timeout_sec="720"
sandbox_concurrency="250"
cli_stable_version=""
cli_beta_version=""
cli_next_version=""
fi

suite_json="$(jq -Rc 'split(",") | map(gsub("^\\s+|\\s+$"; "")) | map(select(length > 0))' <<< "$suite")"
Expand Down Expand Up @@ -177,9 +153,6 @@ jobs:
echo "sandbox_concurrency=$sandbox_concurrency"
echo "filter_changed=$filter_changed"
echo "do_merge=$do_merge"
echo "cli_stable_version=$cli_stable_version"
echo "cli_beta_version=$cli_beta_version"
echo "cli_next_version=$cli_next_version"
} >> "$GITHUB_OUTPUT"

- name: Checkout
Expand Down Expand Up @@ -319,9 +292,6 @@ jobs:
VERCEL_TOKEN: ${{ secrets.VERCEL_TOKEN }}
VERCEL_TEAM_ID: ${{ secrets.VERCEL_TEAM_ID }}
VERCEL_PROJECT_ID: ${{ secrets.VERCEL_PROJECT_ID }}
SUPABASE_CLI_STABLE_VERSION: ${{ needs.prepare.outputs.cli_stable_version }}
SUPABASE_CLI_BETA_VERSION: ${{ needs.prepare.outputs.cli_beta_version }}
SUPABASE_CLI_NEXT_VERSION: ${{ needs.prepare.outputs.cli_next_version }}
steps:
- name: Checkout
uses: actions/checkout@9f698171ed81b15d1823a05fc7211befd50c8ae0 # v6.0.3
Expand Down Expand Up @@ -350,9 +320,6 @@ jobs:
echo "OPENAI_API_KEY=${OPENAI_API_KEY}"
echo "AI_GATEWAY_API_KEY=${AI_GATEWAY_API_KEY}"
echo "XAI_API_KEY=${XAI_API_KEY}"
echo "SUPABASE_CLI_STABLE_VERSION=${SUPABASE_CLI_STABLE_VERSION}"
echo "SUPABASE_CLI_BETA_VERSION=${SUPABASE_CLI_BETA_VERSION}"
echo "SUPABASE_CLI_NEXT_VERSION=${SUPABASE_CLI_NEXT_VERSION}"
} > .env

- name: Run evals
Expand Down
4 changes: 2 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ Common workflows:

## CLI evals

The CLI team owns `evals/cli/` and its results. CLI evals run on `codex-gpt-6-luna-cli-{pinned,stable,beta,next,nodaemon,absent}` under `experiments/cli/`: `pinned` runs the repo's pinned CLI version, `stable`/`beta`/`next` install the latest stable, beta or next CLI, and `nodaemon`/`absent` additionally force Docker-less sandboxes — comparing the same scenario across CLI environments.
The CLI team owns `evals/cli/` and its results. CLI evals run on `codex-gpt-6-luna-cli-{pinned,stable,beta,next,nodaemon,absent}` under `experiments/cli/`: `pinned` runs the repo's pinned CLI version, `stable`/`beta`/`next` install the CLI that npm's `latest`, `beta` or `next` dist-tag points at, and `nodaemon`/`absent` additionally force Docker-less sandboxes — comparing the same scenario across CLI environments.

Which evals each arm picks up:

Expand All @@ -94,6 +94,6 @@ Which evals each arm picks up:
Common workflows:

- **Add or change a CLI eval.** Add the scenario under `evals/cli/<id>/` (see [Adding an eval](#adding-an-eval)); set `needsDocker: false` in its `PROMPT.md` frontmatter if it can run without Docker, open a PR, and add the `run-evals-changed` label. Results for the changed evals are committed back to your branch and viewable in the Vercel preview.
- **Refresh every CLI eval.** Dispatch the [Refresh eval results](https://github.com/supabase/evals/actions/workflows/eval-refresh.yml) workflow on `main` with `suite: cli` and `experiment_suite: cli`. It opens a draft PR with the updated `cli-eval-results.json` for you to review and merge. Leave `cli_stable_version`/`cli_beta_version`/`cli_next_version` blank to resolve npm's latest dist-tags, or pin them to reproduce a specific run.
- **Refresh every CLI eval.** Dispatch the [Refresh eval results](https://github.com/supabase/evals/actions/workflows/eval-refresh.yml) workflow on `main` with `suite: cli` and `experiment_suite: cli`. It opens a draft PR with the updated `cli-eval-results.json` for you to review and merge. `cliVersion` in an experiment's `localStackRuntime` takes anything `npm view supabase@<spec>` resolves (an exact version, a dist-tag like `latest`/`beta`/`next`, or a range like `^2.120.0`), resolved once per run; a spec that fails to resolve fails only the pairs that use it. To reproduce a specific run, set an exact version.
- **Analyze results over time.** Every merge that changes `cli-eval-results.json` appends a snapshot to [`cli-results.jsonl`](https://supabase.github.io/evals/cli-results.jsonl) on GitHub Pages, alongside the [benchmark](https://supabase.github.io/evals/results.jsonl), [regression](https://supabase.github.io/evals/regression-results.jsonl), and [docs](https://supabase.github.io/evals/docs-results.jsonl) histories.
- **Run the unit tests.** `pnpm --filter @supabase-evals/framework test:cli-lib` (the CLI skip predicates in `experiments/cli/lib/` plus every CLI eval's scorer tests) — also part of `pnpm test`.
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ An eval's optional `local/` directory is copied into the sandbox workspace befor

Set `cliVersion: 2.109.1` in an eval's frontmatter when it requires a specific Supabase CLI release. This overrides an experiment's `localStackRuntime({ cliVersion })` setting; otherwise the runtime setting or repository-wide default applies.

An experiment can instead pass `localStackRuntime({ cliVersion: 'stable' })`, `'beta'` or `'next'` to track npm's dist-tag for the `supabase` package rather than an exact version, resolved once at session start. An eval's own `cliVersion:` pin still wins over either form.
An experiment's `localStackRuntime({ cliVersion })` takes anything `npm view supabase@<spec>` resolves: an exact version, a dist-tag such as `'latest'`, `'beta'` or `'next'`, or a semver range like `'^2.120.0'`. It is resolved once per run, and a spec that fails to resolve fails only the pairs that use it. An eval's own `cliVersion:` pin still wins over any of these.

An experiment can pass `localStackRuntime({ docker: 'no-daemon' })` or `'absent'` to stage a sandbox where the Docker daemon is unreachable or the `docker` binary is missing entirely, instead of the default `'available'`. `needsDocker` defaults to `true`; set it `false` in an eval's frontmatter when the scenario can run, and is meaningful, without a Docker daemon (e.g. starting the stack is the agent's own job) — that's what lets a Docker-less experiment pick the eval up.

Expand Down
Loading