Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
158 commits
Select commit Hold shift + click to select a range
8044301
Remove Copilot on Rails feature
Jun 8, 2026
70dd310
Revert "Remove Copilot on Rails feature"
Jun 8, 2026
a59a761
Update Copilot on Rails views (#1449)
motm32 Jun 8, 2026
b9eaca9
Add `azure-debug-plan` custom agent and instructions (#1475)
MicroFish91 Jun 8, 2026
be42ab1
Add `azure-debug-generate` custom agent and instructions (#1478)
MicroFish91 Jun 8, 2026
4605511
Move skills to resources/agents instruction files (#1482)
nturinski Jun 8, 2026
bb92297
Improvements to dependency instructions for `azure-debug-generate` (#…
MicroFish91 Jun 9, 2026
a593097
Allow executing debug configurations in the project view (#1485)
MicroFish91 Jun 12, 2026
12b3b12
Improvements to compound launch/tasks & make sure TypeScript source m…
MicroFish91 Jun 15, 2026
9da8efa
add ui preview to planning doc (#1484)
nturinski Jun 15, 2026
4355450
Add command to download agent instruction files to workspace (#1488)
nturinski Jun 15, 2026
a151b0d
Add activation events for our extension when we detect CoR artifacts …
nturinski Jun 16, 2026
96eeffe
Change instructions to rename src folder to services for monorepos (#…
nturinski Jun 16, 2026
8e50ce2
Give folder names more service specific names (#1499)
nturinski Jun 17, 2026
f24654a
Improvements to prerequisite finding for `azure-debug-plan` (#1505)
MicroFish91 Jun 17, 2026
501d137
Autopilot Mode (#1497)
nturinski Jun 18, 2026
729299d
Make changes for accessibility and add loading view (#1494)
motm32 Jun 19, 2026
3910d84
Add fixes for next steps views to work properly (#1510)
motm32 Jun 22, 2026
48a5800
Forbid creating anything that is not app code (#1500)
nturinski Jun 23, 2026
f271d0a
Add prerequisite refresh (#1512)
motm32 Jun 23, 2026
205a334
Use a more service-centric planning paradigm (#1508)
nturinski Jun 23, 2026
f08ac6d
Fix view issues (#1513)
motm32 Jun 23, 2026
a432d9c
Make prerequisites logic shared, add a prerequisites view to the scaf…
MicroFish91 Jun 23, 2026
feae13c
Improve loading responsiveness of preview UI and add ability to write…
MicroFish91 Jun 23, 2026
6ccc4af
Add azure-project-integrate agent instructions (#1514)
nturinski Jun 23, 2026
ed8500f
Update instructions to properly capture multiple data stores on requi…
MicroFish91 Jun 23, 2026
4f7a2f6
Nat/integrate integration (#1517)
nturinski Jun 23, 2026
b881197
Fix comments showing up in the debug plan webview (#1520)
MicroFish91 Jun 23, 2026
61f52f3
add fix (#1521)
motm32 Jun 24, 2026
b0b4b51
Add refresh to scaffold prerequisites view (#1522)
motm32 Jun 24, 2026
b0ea35e
Fix feedback drawer issues (#1523)
motm32 Jun 24, 2026
9c26061
Multiple improvements to prerequisites listing and installation warni…
MicroFish91 Jun 25, 2026
6b737d2
Improve prerequisites detection (#1526)
MicroFish91 Jun 25, 2026
6e000d5
Fix lint for build (#1527)
MicroFish91 Jun 25, 2026
9679e7e
Add a warning about using python/.net (#1524)
nturinski Jun 30, 2026
44c8f4a
Increase max requests when using copilot on rails flow (#1531)
nturinski Jun 30, 2026
a634887
Change warning labels and remove some verbose comments (#1534)
nturinski Jun 30, 2026
94c698e
UI changes from feedback (#1529)
motm32 Jun 30, 2026
fab8c70
Greatly improve performance of requirements.json. (#1542)
nturinski Jul 8, 2026
47959c4
More UI fixes (#1541)
motm32 Jul 8, 2026
09c920e
Add "Create with Copilot" button when no folder is opened (#1544)
motm32 Jul 10, 2026
39e1613
Resume session (#1543)
nturinski Jul 14, 2026
f1d22c8
Increase max width (#1548)
motm32 Jul 15, 2026
b990c5a
Improve the UX preview plan project (#1545)
nturinski Jul 15, 2026
2eb35c2
CoR: Cherry pick MCP support, add CoR commands as MCP, wrap commands …
MicroFish91 Jul 16, 2026
9ee7255
Make it so autopilot carries over to new instances (#1539)
motm32 Jul 16, 2026
d0505fa
Add loading view for Running API tests (#1553)
motm32 Jul 16, 2026
6dec530
Requested small UI changes (#1551)
motm32 Jul 17, 2026
fef3b15
CoR: Add browser detection as part of identifying prerequisites and a…
MicroFish91 Jul 17, 2026
b3ed02f
Add a "Report issue" button to the loading views (#1559)
motm32 Jul 20, 2026
628fd83
Add a "Need help" button to loading views so users can easily resume …
motm32 Jul 20, 2026
14daf4c
CoR: Record local diagnostics metadata for originating prompt and cre…
MicroFish91 Jul 20, 2026
3605877
Add warning for Azure resources emulators with limited support (#1554)
motm32 Jul 22, 2026
a2b204a
Add telemetry for refreshing prerequisites (#1567)
motm32 Jul 22, 2026
b673f13
Add no-datastore project requirement (#1572)
alexweininger Jul 22, 2026
b365edc
Wait for custom agents before launch (#1571)
alexweininger Jul 22, 2026
7fe3664
Add a model selector to the landing page (#1573)
motm32 Jul 23, 2026
c2c9e48
Add telemetry for next steps buttons (#1566)
motm32 Jul 23, 2026
968131f
Improve deployment plan parsing (#1574)
alexweininger Jul 23, 2026
48df8ea
Open next steps view again after API tests are run (#1595)
motm32 Jul 24, 2026
4b70b20
Add telemetry and diagnostics for interactions with the local debug p…
MicroFish91 Jul 24, 2026
ecbff45
Add telemetry and diagnostics for interactions with the project scaff…
MicroFish91 Jul 24, 2026
8cbdcf6
Add telemetry and diagnostic for interactions with the requirements v…
motm32 Jul 24, 2026
8179c86
Add telemetry and diagnostics for interactions with the deployment pl…
MicroFish91 Jul 24, 2026
62e7700
Add some missing plan telemetry points (#1602)
MicroFish91 Jul 26, 2026
e2ce62c
CoR: Record ISO timestamps (#1605)
MicroFish91 Jul 27, 2026
74356cb
CoR: Record basic system / model specs for telemetry and diagnostics …
MicroFish91 Jul 27, 2026
3621f2e
CoR: Standardize command and telemetry ids (#1603)
MicroFish91 Jul 27, 2026
5cd76a9
CoR: Delete some dead code (#1612)
MicroFish91 Jul 27, 2026
bb9bb54
Add a warning for when the GHC4A extension/skills are not installed w…
motm32 Jul 28, 2026
9cd0225
Record autopilot telemetry automatically for all CoR commands (#1613)
MicroFish91 Jul 28, 2026
05b883d
CoR: Improve prerequisites detection to be more shell agnostic (#1558)
MicroFish91 Jul 28, 2026
f69cab0
CoR: Add extra telemetry for `openWithChatAgent` family of functions …
MicroFish91 Jul 28, 2026
3e2d1e9
Change agent to always opening the requirements view (#1621)
motm32 Jul 28, 2026
d0574eb
CoR: Add a way to report issue to GitHub with attached diagnostics in…
MicroFish91 Jul 28, 2026
0524202
Write the status for approved project plan (#1619)
motm32 Jul 28, 2026
14da9d0
Fix #1624: guard azd hook schema in deploy agent to stop deploy retry…
alexweininger Jul 29, 2026
d48fffa
Fix JSON formatting in settings.json
nturinski Jul 29, 2026
8d0c7b8
CoR: Remove `onStartupActivation` by making `Azure Project` view alwa…
MicroFish91 Jul 30, 2026
c4c4786
Harden scaffold instructions so the Approve-UI preview iframe reliaby…
nturinski Jul 30, 2026
6dc06fa
Merge from main (#1630)
nturinski Jul 30, 2026
7ba9115
Merge branch 'main' into feat/CoR
Jul 30, 2026
6372547
CoR: Separate system info recording behavior for diagnostics vs. tele…
MicroFish91 Jul 30, 2026
ffeb780
add deploy-ready prebuilt artifact contract (#1643)
nturinski Aug 3, 2026
a867edc
CoR: Invoke extension commands through the VS Code API (#1640)
MicroFish91 Aug 3, 2026
658d8d2
Add frontend-to-backend deployment readiness contract across plan/int…
nturinski Aug 5, 2026
9a9defb
Potential fix for pull request finding
nturinski Aug 5, 2026
3424e21
Make live-client API base framework-agnostic (not Vite-specific)
nturinski Aug 5, 2026
1d949aa
Illustrate distinct Production Source values in plan example (Managed…
nturinski Aug 5, 2026
9c2938a
CoR: Record debug plan approval telemetry even when autopilot is acti…
MicroFish91 Aug 6, 2026
9961aa2
docs: add Create New Project with Copilot guide and support runbook (…
nturinski Aug 13, 2026
8b26ba6
CoR: Add recent-prompt history navigation for prompts on the landing …
MicroFish91 Aug 14, 2026
070f507
CoR: More improvements to local debug agent instructions (#1663)
MicroFish91 Aug 14, 2026
84265ac
Use the app onboarding pipeline (#1665)
nturinski Aug 14, 2026
f794593
scaffold: document monorepo self-contained deploy; bundling optional …
nturinski Aug 17, 2026
d59a797
CoR: Add debug session watcher for project debug telemetry (#1666)
nturinski Aug 17, 2026
936082f
Fix docker-compose credential corruption and blank local database cre…
alexweininger Aug 18, 2026
d9e058a
Merge branch 'main' into feat/CoR
Copilot Aug 20, 2026
298f04a
Fix lint errors in azure-deploy agent reference schema files
Copilot Aug 20, 2026
b01c7ae
CoR: Remove unnecessary prompt input property from some mcp tools (#1…
MicroFish91 Aug 20, 2026
398f7ed
CoR: Add `debugAnyway` as a default workspace setting (#1677)
MicroFish91 Aug 21, 2026
d93a9c9
CoR: In autopilot, reuse project-plan prerequisites for the debug pla…
MicroFish91 Aug 21, 2026
c586821
CoR: Ensure we are using the local harness during runs (#1680)
MicroFish91 Aug 22, 2026
309b01f
CoR: Resolve workspace trust errors that block GitHub Copilot Chat fr…
MicroFish91 Aug 24, 2026
4df6972
Add a deployment results view (#1687)
motm32 Aug 24, 2026
566e7f4
Add vally framework and graders for azure-project-plan (#1683)
motm32 Aug 25, 2026
d6e18de
Run the Vally project-plan eval on MSBench against a real extension b…
alexweininger Aug 25, 2026
55c2dda
Run the real Vally validators in MSBench as exec: assertions (#1695)
alexweininger Aug 25, 2026
2b1e646
Convert the eval scripts to TypeScript (#1696)
alexweininger Aug 25, 2026
e4e719f
Stop run.sh from reporting throttled and raced runs as results (#1698)
alexweininger Aug 26, 2026
28ac4d5
Raise the MSBench run timeouts so long end-to-end runs aren't killed …
alexweininger Aug 26, 2026
9a2373a
Fix the evals lint break left by the TypeScript conversion (#1705)
alexweininger Aug 26, 2026
3e47c9d
Verify what actually ran before trusting a run's results (#1701)
alexweininger Aug 26, 2026
a7449ca
Fail eval runs that die early instead of giving them partial credit (…
alexweininger Aug 26, 2026
49930fa
Add `evals/msbench/regrade.ts`: re-grade past MSBench runs for zero t…
alexweininger Aug 26, 2026
5d4e031
Add the scaffold-phase eval gates from #1693 (#1707)
alexweininger Aug 26, 2026
258482e
Add the local-debug eval gates from #1694 (#1708)
alexweininger Aug 26, 2026
bd0ed8f
Turn on MSBench run queueing, pin smoke mode off, and document the ar…
alexweininger Aug 26, 2026
9dd6d8f
Remove .vally.yaml paths that point at directories that never existed…
alexweininger Aug 26, 2026
13d1561
Delete the retired headless eval runner (#1709)
alexweininger Aug 26, 2026
cabaa4e
Measure what a real multi-turn chain costs, with the first multi-turn…
alexweininger Aug 26, 2026
ad41024
Count the sub-agent trajectories: reported uncached tokens were 2.5x …
alexweininger Aug 26, 2026
a41d578
Split the MSBench config into base, phase and stimulus layers (#1715)
alexweininger Aug 26, 2026
551e0e8
Add fidelity gates: did the agent build what it planned? (#1721)
alexweininger Aug 26, 2026
0d844fb
Audit the gates themselves, not just the product (#1718)
alexweininger Aug 26, 2026
fd85354
Add the runtime gates: does the generated project actually run? (#1719)
alexweininger Aug 26, 2026
c8ea2cf
Add a debug-breakpoint eval gate that hits a real breakpoint (#1717)
alexweininger Aug 26, 2026
5853938
Stop the grader scanner reading English as code (#1727)
alexweininger Aug 26, 2026
8f690c4
Refuse to audit a run MSBench has not marked complete (#1725)
alexweininger Aug 26, 2026
d196494
Make a project type a data file: the stack schema (#1726)
alexweininger Aug 26, 2026
fafab22
Let the health gate fail for the thing it exists to check (#1728)
alexweininger Aug 26, 2026
ed16387
Stop reporting "my parser found nothing" as "there is nothing to chec…
alexweininger Aug 26, 2026
7c0d331
Derive which gates a stack runs, instead of hand-wiring them (#1731)
alexweininger Aug 26, 2026
32c16ba
Flag gates that never fail and sometimes decline to answer (#1729)
alexweininger Aug 26, 2026
f3153fb
Record the part of a gate that nothing tests (#1732)
alexweininger Aug 26, 2026
ddbd1e7
Make "run.sh works on a clean machine" a check, not a comment (#1733)
alexweininger Aug 26, 2026
9707d01
Read the stack declaration, and stop asserting that a worker listens …
alexweininger Aug 26, 2026
b4603a8
Stop two runtime gates requiring the declaration they consume (#1736)
alexweininger Aug 26, 2026
03ce495
Tell gate-health which reds we already agreed to pay for (#1735)
alexweininger Aug 26, 2026
0f8d171
Stop a typo in the gate table from silently unwiring a gate (#1737)
alexweininger Aug 26, 2026
a199d50
Make the unexplained gate rows visible rather than clean (#1738)
alexweininger Aug 26, 2026
3d3336f
Stop the README implying the gate table is verified (#1739)
alexweininger Aug 26, 2026
67e65f0
Stop demanding a Database section from a project that has no datastor…
alexweininger Aug 26, 2026
cdf15ba
Wire the eight merged scaffold and debug graders to real MSBench stim…
alexweininger Aug 26, 2026
dcad7cb
Make the deny-list the default way to say "this fact must be somethin…
alexweininger Aug 26, 2026
35a7070
Wire the debug probe into MSBench, plus one cheap run to prove it ins…
alexweininger Aug 26, 2026
fc75c88
Stop gate-health rebuilding the declared-gap key by hand (#1742)
alexweininger Aug 26, 2026
c735167
Fix five places the MSBench docs and config disagreed with themselves…
alexweininger Aug 26, 2026
fe6fd24
Fix the two things that were only ever wrong on the machine that matt…
alexweininger Aug 26, 2026
5c29c51
CoR:  Add first-launch browser debugging troubleshooting guidance for…
MicroFish91 Aug 26, 2026
4599961
CoR: Improve `azure-debug-generate` troubleshooting docs (#1745)
MicroFish91 Aug 26, 2026
f71349c
Probe the one mechanism a phase chain can still use, before building …
alexweininger Aug 26, 2026
3318bc1
Report a gate that never ran as not-attempted, not as a failure (#1747)
alexweininger Aug 26, 2026
9358472
CoR: Fix webview syntax error from apostrophes in initial data (#1748)
nturinski Aug 26, 2026
a872387
Fail the job when the assertions failed, instead of reporting green (…
alexweininger Aug 26, 2026
c635a49
Inventory ARM resources before and after deployment to give a clean u…
nturinski Aug 26, 2026
70f5f58
Merge remote-tracking branch 'origin/feat/CoR' into nat/deploy-readin…
Copilot Aug 26, 2026
7147b27
Resolve merge conflicts with feat/CoR
Copilot Aug 26, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
73 changes: 73 additions & 0 deletions .github/instructions/copilot-on-rails-docs.instructions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
---
description: 'Keep the Create New Project with Copilot (Copilot on Rails) user guide and support runbook in sync whenever the feature changes.'
applyTo: "src/webviews/copilotOnRails/**, src/commands/copilotOnRails/**, src/chat/tools/copilotOnRails/**, src/utils/copilotOnRails/**, src/tree/project/**, resources/agents/**"
---

# Keep the Copilot on Rails docs in sync

You are editing the **Create New Project with Copilot** feature (codename *Copilot on Rails*, command prefix
`copilotOnRails.`). Its end-user guide and support/triage runbook lives at
[docs/copilot-create-project.md](../../docs/copilot-create-project.md).

**Rule:** any change to this feature's user-visible behavior, surfaces, or support flow must be reflected in
that document **in the same change**. Treat the doc as part of the feature — a change that alters behavior
without updating it is incomplete.

## When a change requires a doc update

Update the matching section of `docs/copilot-create-project.md` when you:

| Change | Section(s) to update |
| --- | --- |
| Add / rename / remove a `copilotOnRails.*` command (TS handler **or** `package.json` / `package.nls.json`) | UI surfaces reference (Part 3), Commands appendix, and the relevant stage |
| Add / rename / remove an MCP tool (`src/chat/tools/copilotOnRails/**`) | The MCP tools table and the pipeline diagram |
| Add / change / remove a webview or its behavior (`src/webviews/copilotOnRails/**`) | UI surfaces table, the stage that uses it, and its screenshot |
| Add / change / remove an agent or a hand-off (`resources/agents/**`) | The agents table, the Mermaid pipeline diagram, and the affected stage |
| Change a `.azure/*` artifact, the `.github/agents` download behavior, or a `workspaceState` key | Files & state |
| Change what diagnostics capture, or the Report Issue / Inspect Diagnostics behavior | Support & triage runbook, including the "What the diagnostics contain (privacy)" section |
| Change the launch / resume / empty-folder / autopilot flow | Launching, Resuming a session, and Autopilot mode |

## New or changed UI — flag screenshots to re-capture

Screenshots are captured by hand and stored separately, so the agent can't re-shoot them. When your change
touches the UI, **tell the developer which images to refresh** and why:

- **Altered an existing screen** (relabeled or moved control, restyled view, new or removed field, changed
copy, different states): its screenshot is now **stale even though the placeholder already exists**. Name
the affected file(s) and say in one line what changed.
- **Added a brand-new screen or state**: add the matching `📷` placeholder (the blockquote plus its centered
`<p align="center"><img …></p>` reference — images in this doc are centered, not raw `![]()`) and a
screenshot references section, then tell the dev it needs a first capture.
- **Removed a screen**: delete its placeholder and checklist entry, and note the removal.
- Never delete an existing placeholder just because its PNG is still missing — the images are captured
separately from the prose.
- When unsure whether a visual change is significant, flag it anyway.

Use this map from source area to the screenshot(s) it backs:

| You changed… | Screenshot(s) to re-capture |
| --- | --- |
| `src/tree/project/**`, the `azureProject` view / welcome content | `01-launch-azure-project-view.png`, `12-azure-project-progress-tree.png` |
| The launch / empty-folder / resume flow (`createProjectWithCopilot.ts`, `resume*`) | `02-empty-folder-prompt.png`, `11-resume-prompt.png` |
| `CreateProjectView` (prompt + model picker) | `03-create-project-prompt.png` |
| `RequirementsView` | `04-requirements-view.png` |
| `ScaffoldPlanView` / plan preview (incl. UI preview cards) | `05-plan-preview.png` |
| `FrontendPreviewView` (Approve UI) | `06-frontend-preview-approve-ui.png` |
| `ScaffoldNextStepsView` | `07-scaffold-next-steps.png` |
| `LocalPlanView` (debug plan) | `08-debug-plan-view.png` |
| `LocalDevNextStepsView` | `09-debug-next-steps.png` |
| `DeploymentPlanView` | `10-deployment-plan-view.png` |
| `DeployResultView` | `15-deployment-results-view.png` |
| `reportIssue` (issue template) | `13-report-issue-github.png` |
| `inspectDiagnostics` (JSON payload) | `14-inspect-diagnostics-json.png` |

## Before you finish

- **Report screenshots to the developer:** in your summary, list every image your change makes stale (by
filename, with a one-line reason) plus any placeholders you added or removed, so they can capture or
refresh them. If your change touched no UI, say so.
- Re-read the affected sections and confirm every command id, MCP tool name, agent name, file path, and
view→command mapping still matches the code you changed.
- Keep the reference tables and the Mermaid pipeline diagram accurate.
- If nothing user-visible changed (a pure internal refactor), no doc update is needed — note that briefly
instead of editing the doc.
110 changes: 110 additions & 0 deletions .github/workflows/agent-contracts.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
name: Agent Contracts

# The half of the Vally eval suite that needs no Copilot credentials.
#
# Running the agent itself now happens on MSBench (see msbench-evals.yml), but
# the checks below still earn their place in PR CI: they are fast, they need no
# token, and they guard the graders and agent assets that the MSBench run
# depends on. A broken grader or a drifted instruction file would otherwise only
# surface as a confusing eval failure much later.
#
# Deliberately absent: anything that drives a live agent. That now happens on
# MSBench (see msbench-evals.yml), and the files that did it headlessly —
# run-eval.cjs, check-copilot-auth.ts, check-gate-tools.ts and the SDK executor —
# have been deleted rather than left unreferenced in the tree.
on:
pull_request:
paths:
- 'evals/**'
- 'resources/agents/**'
- '.vally.yaml'
- '.github/workflows/agent-contracts.yml'
workflow_dispatch:

env:
NODE_VERSION: '22'

jobs:
contracts:
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: read
steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}

- name: Install eval dependencies
run: npm ci
working-directory: evals

- name: Add Vally CLI to PATH
run: echo "$GITHUB_WORKSPACE/evals/node_modules/.bin" >> "$GITHUB_PATH"

# A rule removed from the shipped agent should not linger in the eval's
# copy of it, or the eval grades a prompt we no longer ship.
- name: Check agent instruction drift
run: npm run drift
working-directory: evals

# Graders run straight off TypeScript source, so a type error is a broken
# grader — catch it before it costs a full eval run.
- name: Type-check evals
run: npm run typecheck
working-directory: evals

# The grader import scanner decides what gets staged into the container, and a
# naive one already failed the build on a prose sentence. These cases are the
# ones it must get right — including that a genuine bare import STILL throws,
# since the guard's entire value is its ability to fail.
- name: Self-test the import scanner
run: npm run imports:self-test
working-directory: evals

# Prove the graders still detect the regressions they claim to detect.
- name: Certify graders
run: npm run certify
working-directory: evals

# Same idea one layer up: prove the stack schema still rejects the sixteen
# broken stack files it claims to reject. A schema that quietly accepts
# everything would let a stack requiring a binary the container does not
# have reach a paid run, which is the cost this check exists to avoid.
- name: Check stack schema
run: npm run stacks:check
working-directory: evals

# A SQL assertion has no grader filename, so its `comment` is the only
# stable identity it carries into stored run results. Rewording one forks
# that gate into a second identity with no history — silently, since the
# assertion still runs and still passes. The liveness sentinel reached ten
# different wordings across eleven stimuli before anyone noticed, and no
# identity scheme can repair that retroactively. Comments are the one place
# paraphrasing is normally harmless, so nothing but a check will stop it.
- name: Check shared assertion comments are canonical
run: npm run gates
working-directory: evals

# Run through the evals package so the spec is actually linted: a bare
# `vally lint` at the repo root discovers no skills and silently passes.
- name: Lint eval specs
run: npm run lint
working-directory: evals

# `run.sh` promises a clean machine can execute it, and the MSBench eval job
# depends on that: it installs nothing before running `run.sh --skip-build`.
# A script it invokes that statically imports a package therefore fails on
# the one path where failure costs money. That rule was written down, in
# prose, and then broken by the next change to the same file family — so it
# is a mechanism now rather than a hope.
#
# Deliberately last: this step moves `evals/node_modules` aside and restores
# it in a `finally`. Running it after everything else means that even a
# catastrophic failure to restore cannot make an unrelated step fail with a
# confusing error.
- name: Check run.sh works on a clean machine
run: npm run clean-machine:check
working-directory: evals
150 changes: 150 additions & 0 deletions .github/workflows/msbench-evals.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,150 @@
name: MSBench Evals

# Runs the project-plan eval on MSBench against a real VSIX build of this
# extension, rather than against agent instructions in isolation.
#
# Split into two jobs on purpose:
#
# build - packages the VSIX and asserts what ended up inside it, using the same
# guards run.sh applies locally. The shared build template
# (microsoft/vscode-azuretools jobs.yml) already compiles and packages,
# but it never inspects the archive, so a .vscodeignore rule that drops
# resources/agents/ passes every other check in the repo and then shows
# up as an agent that mysteriously ignores its instructions. Needs no
# credentials, so it catches that on a PR instead of in an MSBench run.
# eval - submits to MSBench. Needs Azure auth, so it is manual-only; see
# evals/msbench/README.md ("Running in CI") for the one-time setup.
#
# Unlike the credential-free gates in agent-contracts.yml, this cannot ride on
# GITHUB_TOKEN: MSBench runs on CES, which authenticates callers by Entra client id.
on:
pull_request:
paths:
- 'evals/msbench/**'
- '.github/workflows/msbench-evals.yml'
# The build job's lasting value is asserting what ends up *inside* the
# VSIX, which the shared build template does not check. Both inputs to
# that live outside evals/, so they have to trigger it themselves.
- '.vscodeignore'
- 'resources/agents/**'
workflow_dispatch:
inputs:
benchmark:
description: 'Benchmark instance to borrow for its container image'
required: false
default: 'vscbench.say_hello'
stimulus:
description: 'Stimulus to submit. The default is the cheapest one whose answer we already know.'
required: false
default: 'scaffold-unapproved-plan'

env:
NODE_VERSION: '22'
PYTHON_VERSION: '3.12'

jobs:
# Everything that can be verified without credentials.
build:
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}

# `stage-graders.ts` copies an allowlist of dependency-free packages out of
# evals/node_modules into the staged tree, so a grader's bare import (today
# `jsonc-parser`, for the JSON-with-comments in launch.json) resolves inside
# the container, which has no install step. It hard-errors rather than
# staging a partial tree, so without this the build job fails before the
# VSIX assertions it exists to run.
- name: Install eval dependencies
run: npm ci
working-directory: evals

- name: Build and verify the VSIX
run: ./evals/msbench/run.sh --build-only

- uses: actions/upload-artifact@v4
with:
name: msbench-vsix
path: evals/msbench/assets/extensions/*.vsix
retention-days: 7

eval:
# Manual only: the Azure identity has to be allowlisted by the MSBench team
# before this can pass, so running it on PRs would only ever be red.
if: github.event_name == 'workflow_dispatch'
needs: build
runs-on: ubuntu-latest
# A cold run is ~15 min; the timeout is generous so a slow queue does not
# look like a product failure.
timeout-minutes: 60
permissions:
contents: read
id-token: write # Fetch an OIDC token for azure/login.
steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}

- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHON_VERSION }}

- uses: actions/download-artifact@v4
with:
name: msbench-vsix
path: evals/msbench/assets/extensions

# Needed here too, not just in `build`: `--skip-build` skips the VSIX, but
# graders are staged on every invocation because they are read straight off
# the working tree, and staging them needs evals/node_modules for the
# allowlisted packages.
- name: Install eval dependencies
run: npm ci
working-directory: evals

# run.sh mints the MSBench feed token with `az account get-access-token`,
# so it only needs an already-authenticated az. A self-hosted runner that
# is already signed in can skip this step entirely.
- name: Azure login
uses: azure/login@v2
with:
client-id: ${{ secrets.MSBENCH_AZURE_CLIENT_ID }}
tenant-id: ${{ secrets.MSBENCH_AZURE_TENANT_ID }}
subscription-id: ${{ secrets.MSBENCH_AZURE_SUBSCRIPTION_ID }}

# STIMULUS is read from the environment by run.sh, the same way BENCHMARK
# is, so `--stimulus` on the command line still wins for a local run.
#
# It is set explicitly rather than left to run.sh's default, which is
# `photo-app-requirements` — a full planning run whose result would then
# have to be interpreted. The first CI runs are testing the *pipeline*, so
# they use a stimulus whose answer is already known locally (6/6 green):
# a red then means CI is broken, which is the only question being asked.
# A first run against an unknown-answer stimulus cannot separate "CI is
# misconfigured" from "the product changed".
- name: Run the MSBench eval
run: ./evals/msbench/run.sh --skip-build --output "$RUNNER_TEMP/report.json" --data_dir "$RUNNER_TEMP/msbench-data"
env:
BENCHMARK: ${{ inputs.benchmark }}
STIMULUS: ${{ inputs.stimulus }}

# Keep the report and per-instance data so a failure can be diagnosed from
# the transcript and patch rather than by re-running.
- name: Upload run artifacts
if: always()
uses: actions/upload-artifact@v4
with:
name: msbench-results
path: |
${{ runner.temp }}/report.json
${{ runner.temp }}/msbench-data/**
if-no-files-found: ignore
39 changes: 39 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -67,3 +67,42 @@ testWorkspace
test-results.xml
dist
stats.json
results/

# Generated eval skills (built from resources/agents/*.agent.md at run time)
evals/.generated/

# Grader sources staged into the MSBench agent assets by evals/msbench/run.sh.
# Generated from the working tree on every run; a checked-in copy would drift.
evals/msbench/assets/graders/

# Built from evals/msbench/config/{base,stimuli/*}.yaml by run.sh. run-agent.sh
# only ever reads this one filename, so selecting a stimulus means writing it.
evals/msbench/assets/user-overrides.yaml

# The resolved stack, projected as JSON for the container to read (the graders
# run from staged source with no node_modules, so they cannot parse the YAML).
# Written by build-config.ts on every --stack build, alongside the graders rather
# than inside assets/graders/, which stage-graders.ts wipes.
evals/msbench/assets/stack.json

# Starting workspace materialised by evals/msbench/stage-workspace.ts from the
# stimulus's `# seed:` directive. Derived from a checked-in fixture on every run,
# and cleared between stimuli — a checked-in copy would be a third source of
# truth for a document that already has two.
evals/msbench/assets/workspace/

# The known-debuggable project the breakpoint stimulus runs against, copied from
# evals/grader-certification/reference-node-fullstack by run.sh on every run.
# Same reason as the graders above: a second checked-in copy would drift from the
# fixture that evals/debug-probe certifies against, and then a green run here and
# a green certification would stop meaning the same thing.
evals/msbench/assets/fixtures/

# Mutual-exclusion lock taken by run.sh while it owns assets/.
evals/msbench/assets/.run.lock/

# Extraction cache for evals/msbench/regrade.ts, keyed by run id. Downloaded
# from stored MSBench results and reused across iterations so re-grading stays a
# local operation; safe to delete at any time.
evals/msbench/.regrade/
20 changes: 20 additions & 0 deletions .vally.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Vally project configuration
#
# `paths.results` is an *output* directory, created on demand and gitignored, so
# it is correctly absent from a clean checkout. Every other path here is an input
# and must exist — a path that points at nothing lints clean and silently
# contributes no evals, which is worse than a missing-file error.
#
# No `environments.mcpServers` entry. The workflow-tools stand-in existed for the
# headless SDK runner, which has been deleted — the agent now runs on MSBench
# against the extension's real in-process MCP server. The `workflow-tools-*` tool
# names in the specs below are the names Vally matches against a trajectory, and
# do not require a server to be declared here.
paths:
evals: [evals]
results: results

suites:
project-plan:
description: azure-project-plan agent contracts
evals: ["evals/project-plan/**"]
Loading
Loading