Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
c92dd35
feat(evals): add branching entitlement regression evals
claude Oct 7, 2026
d45b12e
chore: refresh eval results [skip ci]
github-actions[bot] Oct 7, 2026
796786c
fix(evals): require a create_branch attempt in 004, classify 003 appr…
Rodriguespn Oct 8, 2026
3f45850
chore(platform-lite): sync advertised OpenAPI entries for org, entitl…
Rodriguespn Oct 8, 2026
9a4c182
refactor: simplify branching eval helpers and platform-lite seeding
Rodriguespn Oct 8, 2026
3ca0ede
refactor(core): move generic MCP tool-call checks into the framework
Rodriguespn Oct 8, 2026
6382546
feat(evals): cost-consent prompt line, upgrade-link judge, TEMP MCP p…
Rodriguespn Oct 8, 2026
ea8f4c2
refactor(evals): shared MCP tool-call checks live in evals/lib
Rodriguespn Oct 8, 2026
9948c53
chore: refresh eval results [skip ci]
github-actions[bot] Oct 8, 2026
899f903
fix(evals): scope shared MCP checks to the Supabase MCP server
Rodriguespn Oct 8, 2026
2e5cc72
chore: refresh eval results [skip ci]
github-actions[bot] Oct 8, 2026
725c432
feat(evals): 002 expects one availability attempt and an upgrade offer
Rodriguespn Oct 8, 2026
6e69094
docs(evals): 003 approval stops after get_cost are a harness limit
Rodriguespn Oct 8, 2026
bd9e67b
docs(evals): get_cost -> confirm_cost is the pre-2026-07-28 cost flow
Rodriguespn Oct 8, 2026
26a63a8
feat(evals): branching evals follow the platform availability design
Rodriguespn Oct 8, 2026
6aa2881
feat(evals): judge only requires saying the plan doesn't include bran…
Rodriguespn Oct 8, 2026
9d8b65a
chore: refresh eval results [skip ci]
github-actions[bot] Oct 8, 2026
29bfc1c
test(platform-lite): /branching answers for a branch's own project ref
Rodriguespn Oct 8, 2026
358afa9
chore: refresh eval results [skip ci]
github-actions[bot] Oct 8, 2026
3063de3
feat(evals): 002/004 pass on one branching-tool unavailable result
Rodriguespn Oct 8, 2026
b491008
chore: refresh eval results [skip ci]
github-actions[bot] Oct 8, 2026
000e909
feat(evals): 002/004 allow parallel first branching calls, fail on re…
Rodriguespn Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,8 @@ Reserve LLM-as-a-judge checks via `ctx.judge()` for semantic or free-form outcom

Prefer building checks declaratively and returning the list in one place instead of accumulating checks within branching logic, so the list remains stable if one path fails.

Reuse shared checks before writing your own. `evals/lib/` holds checks any suite can use, such as counting MCP tool calls (`checkMcpCallCount`) and failing on MCP tool errors (`checkNoMcpToolErrors`). Helpers shared within one suite live in `evals/<suite>/lib/`. Run their tests with `pnpm --filter @supabase-evals/framework test:evals-lib`.

## Adding an experiment

Add a `*.experiment.ts` file under `experiments/<owner>/` for the agent, model, and runtime setup you want to compare. Experiment discovery only scans this owner directory depth, so supporting files can live beside experiments or in nested directories. Reuse the base configs exported from `experiments/presets.ts` where they fit.
Expand Down
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,7 +102,7 @@ Every eval contains:

1. `PROMPT.md` - frontmatter metadata plus the task description the agent sees.
2. `EVAL.ts` - a default-exported scorer.
3. Optional `remote/` - the hosted project's starting state, seeded into platform-lite: `project.sql` (database), `logs.jsonl` (observability logs), `functions/` (already-deployed edge functions).
3. Optional `remote/` - the hosted project's starting state, seeded into platform-lite: `project.sql` (database), `logs.jsonl` (observability logs), `functions/` (already-deployed edge functions), `migrations/` (`<version>_<name>.sql` migration history, applied before `project.sql`), `organization.json` (the org's `name` and `plan`, default `free`).
4. Optional `local/` - the agent's starting files, copied into the sandbox workspace the agent works in (absent means an empty workspace, or no sandbox at all for tools evals).

The two directories mirror Supabase's two environments: `remote/` describes what the customer's hosted project already looks like, `local/` describes what the developer's working directory already looks like.
Expand All @@ -124,6 +124,8 @@ motivation: AI-123

Allowed metadata values are defined in `packages/core/src/eval-metadata.ts`.

Tools evals can also shape the Supabase MCP server: `mcpFeatures: [branching]` enables feature groups on top of the experiment's, and `projectScoped: true` scopes the server to the seeded project (`--project-ref`), which hides account tools like `get_organization` and `get_cost`.

## Eval Modes

There are two runtimes, chosen automatically per eval:
Expand Down
13 changes: 12 additions & 1 deletion apps/framework/harness/run-eval.ts
Original file line number Diff line number Diff line change
Expand Up @@ -148,9 +148,10 @@ function discoverEvals(): EvalManifest[] {
if (!existsSync(root)) return [];
const out: EvalManifest[] = [];
// evals/<suite>/<id>/. The suite folder is what CODEOWNERS scopes by.
// evals/lib/ holds scorer helpers shared across suites.
for (const suiteDir of readdirSync(root)) {
const dir = join(root, suiteDir);
if (!statSync(dir).isDirectory()) continue;
if (suiteDir === 'lib' || !statSync(dir).isDirectory()) continue;
const suite = evalSuiteSchema.parse(suiteDir);
for (const id of readdirSync(dir)) {
const evalDir = join(dir, id);
Expand Down Expand Up @@ -325,13 +326,21 @@ function readSessionSeedArgs(ev: EvalManifest) {
const projectSeedSql = join(ev.remoteDir, 'project.sql');
const logsSeedJsonl = join(ev.remoteDir, 'logs.jsonl');
const functionsSeedDir = join(ev.remoteDir, 'functions');
const organizationSeedJson = join(ev.remoteDir, 'organization.json');
const migrationsSeedDir = join(ev.remoteDir, 'migrations');

return {
projectSeedSql: existsSync(projectSeedSql) ? projectSeedSql : undefined,
logsSeedJsonl: existsSync(logsSeedJsonl) ? logsSeedJsonl : undefined,
functionsSeedDir: existsSync(functionsSeedDir)
? functionsSeedDir
: undefined,
organizationSeedJson: existsSync(organizationSeedJson)
? organizationSeedJson
: undefined,
migrationsSeedDir: existsSync(migrationsSeedDir)
? migrationsSeedDir
: undefined,
pgvector: ev.metadata.product.includes('vectors'),
};
}
Expand Down Expand Up @@ -537,6 +546,8 @@ async function runOne(
await exp.runtime.startSession({
...readSessionSeedArgs(ev),
hostname: agentRunsInSandbox ? '0.0.0.0' : undefined,
projectScoped: ev.metadata.projectScoped,
mcpFeatures: ev.metadata.mcpFeatures,
})
);

Expand Down
3 changes: 2 additions & 1 deletion apps/framework/harness/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,8 @@ export interface EvalManifest {
evalPath: string;
/**
* `remote/` — the hosted project's starting state, seeded into
* platform-lite (project.sql, logs.jsonl, functions/).
* platform-lite (project.sql, migrations/, logs.jsonl, functions/,
* organization.json).
*/
remoteDir: string;
}
3 changes: 2 additions & 1 deletion apps/framework/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
"version": "0.0.1",
"type": "module",
"scripts": {
"check": "pnpm typecheck && pnpm test && pnpm test:framework && pnpm test:cli-lib",
"check": "pnpm typecheck && pnpm test && pnpm test:framework && pnpm test:cli-lib && pnpm test:evals-lib",
"eval": "node --env-file=../../.env --import tsx/esm harness/run-eval.ts",
"eval:upload": "node --import tsx/esm scripts/eval-upload.ts",
"eval:dry": "node --env-file=../../.env --import tsx/esm harness/run-eval.ts --dry",
Expand All @@ -15,6 +15,7 @@
"test:framework": "node --env-file-if-exists=../../.env --import tsx/esm scripts/smoke-framework.ts",
"test:vercel-runner": "vitest run scripts/run-vercel-evals.test.ts scripts/export-results.test.ts lib/cli-args.test.ts lib/experiment-files.test.ts lib/sample-sets.test.ts",
"test:cli-lib": "vitest run --root ../.. --passWithNoTests experiments/cli evals/cli",
"test:evals-lib": "vitest run --root ../.. evals/lib evals/regression/lib",
"export-results": "node --import tsx/esm scripts/export-results.ts",
"demo:mcp": "node --env-file=../../.env --import tsx/esm scripts/mcp-demo.ts",
"demo:executor": "node --env-file=../../.env --import tsx/esm scripts/executor-demo.ts",
Expand Down
Loading