Skip to content

Fix portable test guidance review regressions - #1289

Merged
Evangelink merged 8 commits into
mainfrom
dev/amauryleve/portable-review-corrections
Oct 9, 2026
Merged

Evangelink merged 8 commits into
mainfrom
dev/amauryleve/portable-review-corrections

Conversation

@Evangelink

@Evangelink Evangelink commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Contribute portable TestFx review corrections back to the source repository, after comparing current upstream:

  • Retain explicitly requested Python regressions and required assertions despite production/native-import/tooling blockers; remove only speculative out-of-scope cases.
  • Select .NET discovery from actual runner/command mode, distinguishing positional VSTest, positional bridged MTP and documented named native MTP. Require successful discovery exit and runner-specific output, not masked counting pipelines.
  • Recognize .slnx; preserve explicit .sln/.slnx/.slnf/.csproj; stop on ambiguity/lookup errors; bound Git inventory to requested/containing roots; resolve the exact selected collection graph before classification.
  • Preserve xUnit executable/native runner configuration, reserve TestingPlatformDotnetTestSupport for the VSTest-command-mode bridge, and use mode-aware direct-project/solution verification, retaining classic native runners.
  • Fix executable Pester enum/identity assertions. Complete promised missing-invoice/error/no-update cases in C++, Kotlin, Pester, Ruby and Rust, including Catch2's Approx header. Align plan/test/report: Kotlin9 (6+3 rows), Pester9 (6+3 rows), Ruby10, Rust8; C++9 sections.
  • Preserve assertion packages/imports and correct AwesomeAssertions namespace. Map effective class/assembly Owner traits to method-only MSTest attributes, deduplicate and flag conflicts/shared inheritance.
  • Replace implicit migration commits with phase validation and explicit-user-only commits.
  • Classify xUnit v3 reader rejection. Alternate permitted readers require positively confirmed technical/path-normalization limitations plus authorized access. Stop/report permission/policy/content-exclusion denials without tool/shell/alias/agent bypass or reconstruction. Unknown rejection is unavailable, not missing; disclose incomplete work.
  • Add owning migration fixtures/goldens/authenticated graders for Owner, assertion-library/import/original-assertion preservation and evaluated VSTest state. Remove conflict-prompt answer cues; add an uncued xUnit version-upgrade dormancy guard. Cover complete read-only inventories and all original test assertions, not merely test names/attributes or .Should() presence.
  • Preserve Ruby harness runner failures: use same-process fail-fast Bash, RSpec dry-run summaries or the configured real Minitest task and test summary, never counting pipelines, suppressed stderr or post-failure runner fallback.

Commits: original f3d02293e; first review 5d7609327; examples acf7fb308; runner/routing f5af583de; reader policy eccaebf25; assertion/inventory 82da18db8; Owner assertions aa079b4c89962ca21fdbf471eda78784a1836590; Ruby harness 728658efd75a1270a1d844e42d43f96cf3dbcf2c.

Scope/rollout: Helper-loading/routing, including filter-syntax, overlaps pending #1287 and is untouched. Open #1189 concerns legacy generation tracking/project placement, not these defects. Current test-engineer already has its scratch contract; no deleted generator is restored. TestFx-specific legacy-generator/BOM policies are not portable and remain unchanged upstream. TestFx still pins e468462d8a0900f5278331c1ae12f45e015bb656: this contribution is not yet consumed by that pin. Merged-source refresh and later removal of applicable exceptions remain downstream work. TestFx-specific adaptations to removed vendored migration plugins are not propagated; valid source framework/platform routes remain unchanged.

Related issue

N/A — found during review of microsoft/testfx#11856. No issue closure is implied.

Validation

python -m unittest discover -s tests\dotnet-test\coverage-analysis\graders -p test_setup_discovery.py -v
python -m unittest discover -s tests\dotnet-test\graders -p test_portable_review_examples.py -v

Final exact files: 16/16 discovery tests passed; 17/17 portable guidance tests passed, no skips; actual Pester5.7.1 ran9/9. The parent's original12+7 run and earlier14-test portable run are historical, not the current submitted portable count. Coverage includes entry/ambiguity/lookup/classic classification/root isolation/exact graph; promised invocation/error/non-persistence agreement; three reader rejection dispositions; and three Ruby harness regressions. All three reader guards reject upstream's unconditional fallback plus targeted unsafe mutations (6/6 rejected). These are guidance-contract checks, not simulated access-control enforcement; no denied file was accessed. Subprocesses are timeout-bounded; absent Pester-v5 is explicitly skipped/not executed.

Ruby harness follow-up (728658efd): The portable unittest command above was rerun with a 180s outer subprocess bound: exit 0, 17/17, no skips, including actual Pester9/9. Three new tests check selection/count guidance and execute the two published Bash blocks with synthetic runners: four executions, expected exits 0,0,23,23, preserved stdout summaries/stderr, and exactly one selected runner per invocation (30s bound each). GNU Bash5.3.15(2)-release on Git for Windows. Both unchanged upstream commands were separately reproduced returning Bash exit 0 despite runner exit 23. Eval-quality against origin/main and whitespace check exit 0, existing warnings only. No eval YAML or framework/plugin routes changed in this follow-up. Ruby/Bundler are absent: synthetic exit propagation is not native RSpec/Rake/Minitest integration proof; no tools installed.

Ruby/RSpec and Kotlin/Gradle/JVM tools are absent. rustc --version and cargo --version each exit 2, tool_not_installed; installed launchers are not a usable toolchain. Those three samples have structural evidence only. No follow-up tooling/global settings were installed or changed.

dotnet run --project eng\skill-validator\src\SkillValidator.csproj --no-build -- check --plugin plugins\dotnet-test
dotnet run --project eng\skill-validator\src\SkillValidator.csproj --no-build -- check --plugin plugins\dotnet-test-migration
python eng\eval-quality\check_eval_quality.py --base-ref origin/main
git diff --check
git diff --cached --check
git diff origin/main...HEAD --check

All exit 0: plugin checks 23 skills/10 agents, 6 skills/1 agent, eval-quality with existing warnings only, whitespace clean. Both plugins/regression files were rerun after the reader correction; eval-quality/whitespace repeated after final Owner grading. Validator restore followed a proven missing-assets failure. Owning eval replaces deprecated config: with defaults: and adds capability/risk/journey tags; rubric weights, plugin metadata and external domains are unchanged.

Native/reference contracts: Original xUnit fixtures ran 3/3,3/3,2/2 tests. Migrated goldens ran 3/3,2/2 on MSTest4.4.1, Test.Sdk18.9.0, SDK11.0.100-rc.1.26425.128, net8.0/VSTest. Inputs: xunit.v3 3.2.2, adapter 3.1.5, Test.Sdk 17.13.0. AwesomeAssertions 9.6.0, local/global/project imports and original assertions survive. Owner reflection requires one alice Owner on each of exactly three discoverable methods. Helper digests and complete staged input hashes are authenticated.

dotnet test TestProject.csproj --logger "console;verbosity=normal"
dotnet msbuild TestProject.csproj -getProperty:TargetFramework,EnableMSTestRunner,UseMicrosoftTestingPlatformRunner,TestingPlatformDotnetTestSupport,IsTestingPlatformApplication
dotnet run --project .eval/Contract.csproj
python .eval/assertion_contract.py
python <retained-session-artifacts>\replay_runner_review.py
python <retained-session-artifacts>\replay_migration_review.py --existing
python <retained-session-artifacts>\verify_final_eval_guards.py
python <retained-session-artifacts>\verify_owner_assertions.py

Correct states exit 0, with exact nonzero test counts/markers. Windows replay resolves specification python3 to installed python. Final replay accepts all seven deterministic grader invocations against three reference states. Initial seven mutations reject lost/wrong/duplicate Owner, silent conflict rewriting and separate local/global/project FluentAssertions imports. Runner replay: 17 commands, package VSTest and MSTest.Sdk4.4.1 + UseVSTest=true accepted, SDK references run 3/2 tests, six native/bridged-MTP mutations and a boundary source edit rejected.

Final assertion/inventory replay: 18 commands, arrow/block formatting accepted, four weakened-assertion/comment-decoy mutations rejected, both complete read-only goldens accepted, six add-file mutations rejected (Directory.Build.props, Directory.Build.targets and arbitrary JSON in each case). Inventory covers every file outside known .eval/.git metadata and bin/obj, with exact supplied hashes, not extension filtering.

Final Owner assertion guard: verify_owner_assertions.py exits 0 across nine native command logs. Original three assertion bodies are required alongside reflected names/Owner/discovery metadata. Golden and three individual block-format variants pass. Removing each assertion independently or emptying all three bodies produces four rejected mutations (exit1). The final seven-grader replay, production parser and full-PR eval-quality pass after the authenticated Owner digest update. Logs/scripts/reference fixtures are retained in migration-review-verified, migration-review-final-graders, runner-review, final-eval-guards, owner-assertion-review and final quality logs.

Production probe — executed, non-passing: Installed VallyCLI0.9.0, not claimed equivalent to documented CI0.14:

vally lint plugins\dotnet-test-migration\skills\migrate-xunit-to-mstest -e tests\dotnet-test-migration\migrate-xunit-to-mstest\eval.yaml
vally experiment run <retained-session-artifacts>\production-review\experiment.yaml --dry-run --workers 8
vally experiment run <retained-session-artifacts>\production-review\experiment.yaml --workers 8 --output-dir <retained-session-artifacts>\production-review\results

Lint/dry run exit 0 (one default-scoring warning). Actual run exits 1 after 8/8 executor records completed successfully, zero executor errors/timeouts: four stimuli × baseline/isolated × one run, claude-opus-4.6 executor/judge, normal 8 workers, declared 20m timeout, outer1320s bound. Runtime used f5af583de; subsequent 82da18db8/aa079b4c8 grader strengthening was re-linted and deterministically replayed, not model-rerun.

Authenticated command graders returned 9009 because Windows python3 resolves to an uninstalled Microsoft Store alias. That host prerequisite was classified before content changes. All three positive isolated trials invoked the target; advisory trial did not, but its rubric caught an unwanted MSTest recommendation. The judge also flagged a possible project-import omission, not certified by the blocked deterministic grader. Neither arm is green; no preference W/T/L/powered verdict is claimed. No plugin arm, second family, rerun or broad matrix. Owning suite: 16 preference stimuli + one dormant boundary; four-case probe is deliberately below power floor. Raw results/trajectories/metadata/run.log retained in production-review\results\2026-10-09T13-12-26-559Z, run ID 8ef8bd30-2836-4785-a8cf-b482aa8770b7; Vally cleaned temporary model workspaces.

Real selected collection: Selected three-test project references another project with an intentionally failing test. dotnet test TestProject.csproj --logger "console;verbosity=normal" and dotnet test Selected.slnf --logger "console;verbosity=normal" both exit 0, run only the selected three tests, never the failing dependency. Published inventory agrees. Retained verify_collection_graph.py/collection-graph contain fixture/logs.

Composed runner discovery/execution: Every command exits 0; discovery contains real Invoice_is_discovered; direct/solution runs execute it.

Mode/components Actual fixture commands
VSTest: SDK10.0.401, MSTest3.8.0, Test.Sdk17.13.0 dotnet test Invoice.Tests.csproj --no-build; dotnet test Billing.sln --list-tests --no-build; dotnet test Billing.sln --no-build
Bridged MTP: SDK8.0.425, xunit.v3 3.2.2, MTP1.9.1 dotnet test Invoice.Tests.csproj --no-build; dotnet test Billing.sln --no-build -p:TestingPlatformCaptureOutput=false -- --list-tests; dotnet test Billing.sln --no-build
Native MTP: SDK10.0.401, xunit.v3 3.2.2, MTP1.9.1, no bridge property dotnet test --project Invoice.Tests.csproj --no-build; dotnet test --solution Billing.slnx --list-tests --no-build; dotnet test --solution Billing.slnx --no-build

Qualification: SDK10.0.401 also accepted positional native project syntax (dotnet test Invoice.Tests.csproj --no-build). Guidance aligns documented named forms, not universal parser failure; the initial negative fixture expectation was corrected.

Namespace/attribute probes: AwesomeAssertions9.6.0 with MSTest3.8.0/exact4.4.1. dotnet run -v:q and dotnet run --project <retained-session-artifacts>\native-contracts\metadata\Metadata.csproj -v:q exit 0; invalid class Owner via dotnet build --no-restore -v:q fails CS0592.

C++ native sample: Published blocks/minimal interface header compiled as C++20, MSVC19.51.36260.0, Catch2v3.8.1 commit 2b60af89e23d28eefc081bc930831ee9d45ea58b, after installed vcvars64.bat setup:

cmake.exe -S . -B build -G "NMake Makefiles" -DCATCH2_SOURCE_DIR=<fixture>\Catch2 -DCMAKE_BUILD_TYPE=Debug
cmake.exe --build build --parallel 2
build\invoice_service_tests.exe --reporter console
build\invoice_service_tests.exe --reporter junit --out catch2-results.xml
cmake.exe --build build --target test

All exit 0: 2 cases/9 sections/12 assertions, CTest1/1. NMake ignores parallel; incidental missing-vswhere warning did not prevent build/execution. Missing Catch2 was established before obtaining the pinned dependency. Fixture/log/XML evidence retained.

Original Pester enum failure, .slnx misselection and sibling leakage were reproduced before correction. These are local native/reference checks, not CI results or certified model-quality improvement. No broad Vally/model-family matrix was launched.

Checklist

  • I searched existing issues and pull requests to avoid duplicates.
  • I kept this pull request focused and avoided unrelated refactors.
  • I added or updated tests, evals, or documentation when changing skill or agent behavior.
  • I updated CODEOWNERS when adding or moving owned content. (N/A: existing ownership of both affected plugins/tests covers all paths.)
  • I updated all marketplace manifests when plugin metadata changed. (N/A: no plugin metadata changed.)
  • I updated eng/known-domains.txt for any new external domains referenced by skill content. (N/A: no new domains.)
Evaluation changes only

For every eval-related change, including a new eval:

Three correctness scenarios and one uncued dormancy guard; existing already-MSTest no-op retained. Suite16 preference + one dormancy; four-case probe is not powered quality evidence.

  • Scenarios are necessary, distinct, and use natural prompts.
  • Graders cover the full result, accept the golden result, and reject a realistic mutation.
  • No-op, dormancy, and statistical power are covered where needed. (No-op/dormancy covered; powered quality not measured by this subset.)
  • I ran the applicable production evaluation path and recorded the result above. (Subset completed; exit1, revision/host/quality limitations reported, not a passing verdict.)
  • If this fixes a failed eval, I classified the failure before editing skill content. (Windows prerequisite classified; no prose changed in response to that run.)
  • If this broadly changes routing or behavior, I checked separate model-family evidence. (N/A: no broad routing change or matrix.)

Preserve requested regression coverage, bound coverage discovery to the requested repository, distinguish native test command modes, repair runnable examples, and retain migration metadata and explicit commit authorization. Add focused executable discovery and portable example regression checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 23ac783e-0fae-4968-8b09-efd7ffa72c93
@Evangelink
Evangelink requested a review from a team as a code owner October 9, 2026 11:58
Copilot AI balanced review requested due to automatic review settings October 9, 2026 11:58
@github-actions

github-actions Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
✅ dotnet-test-migration migrate-xunit-to-mstest 1/1 100%
✅ dotnet-test-migration migrate-xunit-to-xunit-v3 15/16 93.8%
ℹ️ dotnet-test code-testing-extensions - N/A (reference-only)
❔ dotnet-test graders — error
Uncovered: dotnet-test-migration/migrate-xunit-to-xunit-v3
  • [CodePattern] sealed (line 178)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Exact-entry discovery can still include unrelated projects, and key migration corrections lack behavioral regression coverage.

3 open findings
What changed in this PR

Updates portable test guidance and adds regression checks for discovery, runner modes, examples, migration metadata, and commit safety.

Changes:

  • Corrects multi-language testing examples and regression-preservation guidance.
  • Improves .NET entry-point, runner-mode, and migration handling.
  • Adds focused Python regression tests.
File Description
tests/​dotnet-test/​graders/​test_portable_review_examples.py Adds guidance regression checks.
tests/​dotnet-test/​coverage-analysis/​graders/​test_setup_discovery.py Tests PowerShell discovery behavior.
plugins/​dotnet-test/​skills/​scaffold-dotnet-test-project/​SKILL.md Clarifies MTP scaffolding and verification.
plugins/​dotnet-test/​skills/​coverage-analysis/​references/​setup-discovery.md Expands entry-point discovery and isolation.
plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​python.md Preserves requested failing regressions.
plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​powershell-examples.md Fixes executable Pester examples.
plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​dotnet.md Makes discovery mode-aware.
plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​cpp-examples.md Completes Catch2 coverage example.
plugins/​dotnet-test-migration/​skills/​migrate-xunit-to-mstest/​references/​mapping-cheatsheet.md Corrects Owner and assertion-library mappings.
plugins/​dotnet-test-migration/​agents/​test-migration.agent.md Replaces implicit commits with validation.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/dotnet-test/graders/test_portable_review_examples.py
Comment thread tests/dotnet-test/graders/test_portable_review_examples.py
@Evangelink
Evangelink enabled auto-merge (squash) October 9, 2026 12:27
@github-actions github-actions Bot added the waiting-on-author PR state label label Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

👋 @Evangelink — this PR has 3 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 12:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/eval.yaml Outdated
Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/eval.yaml
Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/eval.yaml Outdated
Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/eval.yaml
Comment thread tests/dotnet-test/coverage-analysis/graders/test_setup_discovery.py
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 12:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Runner discovery, portable Git handling, assertion grading, and evaluation validation have unresolved defects.

7 open findings
Previously missed (2)

In code that hasn't changed since last review

Medium severity Documented bridge command omits required capture property

plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​dotnet.md:122

The bridged-MTP command omits -p:TestingPlatformCaptureOutput=false, so the bridge's default capture can hide the runner-owned test list that this section requires inspecting. The composed validation in the PR description used that property, so it did not validate the command published here. Include the property (or specify and verify the captured discovery artifact) and update the regression assertion to exercise the exact documented command.

Medium severity Localized Git errors break no-repository detection

plugins/​dotnet-test/​skills/​coverage-analysis/​references/​setup-discovery.md:57

This distinguishes “not in a repository” from a real Git failure by matching Git's English stderr. Git localizes this diagnostic, so the same standalone directory is treated as a fatal lookup error on non-English hosts, contrary to the portable no-repository path below. Force Git's locale to C while capturing this command (restoring the caller's environment afterward), or classify absence without parsing localized text.

🧠 Review effort: Balanced

Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/graders/assertion_contract.py Outdated
Comment thread tests/dotnet-test/graders/test_portable_review_examples.py
Authenticate evaluated VSTest runner state, uncued conflict assessment,
and the xUnit version-upgrade dormancy boundary.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:21
Restrict reader fallback to confirmed authorized technical limitations,
stop on access denials, and disclose unavailable paths and incomplete work.
Guard the three rejection dispositions with focused guidance regressions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread tests/dotnet-test-migration/migrate-xunit-to-mstest/eval.yaml Outdated
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Discovery misses valid modern test-project markers, and two new deterministic graders accept destructive assertion removal.

4 open findings
5 resolved since last review
Previously missed (1)

In code that hasn't changed since last review

Medium severity Recognize established test-project markers

plugins/​dotnet-test/​skills/​coverage-analysis/​references/​setup-discovery.md:120

Both test-project predicates omit established markers such as MSTest.Sdk, Microsoft.Testing.Platform, and <IsTestProject>true</IsTestProject> (compare plugins/dotnet-test/skills/find-untested-sources/SKILL.md:143-148). A neutral-name project using one of those forms is removed from a selected solution graph and can incorrectly produce TEST_PROJECTS:0. Include these markers in both predicates and add a neutral-name MSTest.Sdk regression.

🧠 Review effort: Balanced

Match both original assertion bodies with formatting flexibility and
inventory every non-evaluator project file in read-only migration cases.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:38

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Directory discovery can miss test projects, and two eval contracts do not deterministically reject incomplete or off-target answers.

1 open finding
3 resolved since last review
Previously missed (3)

In code that hasn't changed since last review

Medium severity Non-test project incorrectly limits test discovery

plugins/​dotnet-test/​skills/​coverage-analysis/​references/​setup-discovery.md:96

When $root is a directory, $entry can be the only discovered .csproj even if that project is production code. This branch then limits $projectPaths to that non-test project and never performs the documented requested-directory/containing-Git-root test inventory. For example, requesting src/App when it contains only App.csproj and the repository keeps App.Tests.csproj under tests/ now reports TEST_PROJECTS:0 and stops. Distinguish an explicitly selected test entry from an automatically discovered non-test project, and add this layout to the executable regression.

Medium severity Output checks omit required owner-to-method mappings

tests/​dotnet-test-migration/​migrate-xunit-to-mstest/​eval.yaml:656

These output graders do not prove the owner mappings required by the rubric. A response such as “alice, bob, and charlie conflict; Inherited is shared; manual review is required” passes all four patterns while omitting DirectTests.ConflictingOwners and failing to associate bob/charlie with BillingTests.Inherited/PaymentTests.Inherited. Add deterministic checks tying each conflict to its affected method so this realistic incomplete assessment is rejected.

Medium severity Dormancy output allows conflicting MSTest recommendation

tests/​dotnet-test-migration/​migrate-xunit-to-mstest/​eval.yaml:744

The dormancy contract deterministically checks only that the answer mentions xUnit v3. A response can recommend converting to MSTest as well, pass this output grader and the unchanged-file check, yet violate the routing invariant; the reported bounded run actually exhibited this failure and only the LLM rubric caught it. Add a deterministic output contract that rejects an MSTest conversion recommendation while allowing a concise contrast such as “not MSTest.”

🧠 Review effort: Balanced

Require all three original assertion bodies alongside reflected method
and Owner metadata, with expression/block formatting preserved.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:47

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The bridged-MTP discovery guidance and two evaluation contracts can currently produce incorrect or biased results.

0 open findings

1 resolved since last review
Previously missed (3)

In code that hasn't changed since last review

Medium severity Deterministic grader rejects valid reordered multiline conflict reports

tests/​dotnet-test-migration/​migrate-xunit-to-mstest/​eval.yaml:650

This deterministic grader is order- and line-break-sensitive: a correct response such as Owner conflicts:\n- alice and bob ... fails because owner appears before conflict and . cannot cross the newline. Accept either ordering across lines so valid conflict reports are not rejected.

Medium severity Rubric improperly scores target skill invocation metadata

tests/​dotnet-test-migration/​migrate-xunit-to-mstest/​eval.yaml:749

This rubric scores whether the target skill was invoked, which is harness metadata rather than an answer outcome and biases the pairwise judge toward skill-specific vocabulary. expect_activation: false already enforces dormancy; keep the rubric limited to the xUnit-v3 answer, absence of MSTest conversion, and unchanged files.

Low severity Published bridged command omits output capture disabling property

plugins/​dotnet-test/​skills/​code-testing-extensions/​extensions/​dotnet.md:122

The bridged-MTP command does not expose the test list that the following paragraph requires reviewers to inspect: dotnet test captures the MTP application output unless TestingPlatformCaptureOutput is disabled. The PR's own validated bridged command includes -p:TestingPlatformCaptureOutput=false, but this published command omits it, so it can exit successfully without giving the identities/count needed for the discovery check. Add that property here (or document the runner-owned artifact to inspect) and update the regression assertion accordingly.

🧠 Review effort: Balanced

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-author PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 9, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Copilot AI balanced review requested due to automatic review settings October 9, 2026 15:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Git discovery incorrectly treats localized no-repository diagnostics as fatal lookup failures.

0 open findings

10 resolved since last review
Previously missed (1)

In code that hasn't changed since last review

Medium severity Handle localized Git no-repository diagnostics during root discovery

plugins/​dotnet-test/​skills/​coverage-analysis/​references/​setup-discovery.md:55

This treats only Git's English “not a git repository” diagnostic as the normal no-repository case. Git diagnostics are localized, so the same lookup on a non-English host falls through to Git root lookup failed and aborts discovery instead of continuing with the requested directory. Classify the no-repository state without matching localized stderr (for example, by checking for a containing .git marker while still surfacing broken-marker/lookup failures).

🧠 Review effort: Balanced

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

16 model/target results across 8 targets and 2 models — ✅ 6 improved, ➖ 10 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit aa079b4c89962ca21fdbf471eda78784a1836590; 2 judge models.

Measurement health: 16 expected / 16 observed / 16 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.test-engineer claude-sonnet-5 ✅ Improved n=9; 7W/1T/1L; d=8; p=0.035; net +66.7% — — None.
agent.test-engineer gpt-5.6-luna ✅ Improved n=9; 7W/1T/1L; d=8; p=0.035; net +66.7% — — None.
agent.test-migration claude-sonnet-5 ➖ Improvement signal, unproven n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-migration gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-quality-auditor claude-sonnet-5 ➖ Improvement signal, unproven n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-quality-auditor gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 4W/1T/0L; d=4; p=0.063; net +80.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
coverage-analysis claude-sonnet-5 ➖ Improvement signal, unproven n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded 🟡 0.24 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
coverage-analysis gpt-5.6-luna ➖ Improvement signal, unproven n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded 🟡 0.44 Activation: isolated 7/8; plugin 7/8 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-mstest claude-sonnet-5 ✅ Improved n=16; 11W/3T/2L; d=13; p=0.011; net +56.3%; 1 dormancy excluded 🟡 0.41 — Review overfit evidence.
migrate-xunit-to-mstest gpt-5.6-luna ✅ Improved n=16; 13W/2T/1L; d=14; p=0.001; net +75.0%; 1 dormancy excluded — — None.
migrate-xunit-to-xunit-v3 claude-sonnet-5 ➖ Improvement signal, unproven n=12; 7W/3T/2L; d=9; p=0.090; net +41.7% 🟡 0.28 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-xunit-v3 gpt-5.6-luna ✅ Improved n=12; 11W/1T/0L; d=11; p=0.000; net +91.7% — — None.
scaffold-dotnet-test-project claude-sonnet-5 ➖ Improvement signal, unproven n=10; 5W/3T/2L; d=7; p=0.227; net +30.0% ✅ 0.19 Activation: isolated 9/10; plugin 8/10 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
scaffold-dotnet-test-project gpt-5.6-luna ✅ Improved n=10; 7W/3T/0L; d=7; p=0.008; net +70.0% 🟡 0.48 — Review overfit evidence.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, unproven — agent.test-migration (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +66.7% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 6 paired runs (5W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Plan a staged MSTest v2 to v4 and MTP migration Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Plan a staged MSTest v2 to v4 and MTP migration: Both are directionally correct and satisfy the core requested staging. A is more actionable and safety-oriented, particularly through its explicit baseline stage, regression comparisons, and more detailed gates. Its proposed MTP package wording is somewhat less precise than id...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.test-migration (gpt-5.6-luna)

Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +33.3% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 6 paired runs (3W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline request to write new tests Excluded (activation contract) -100.0% -40.0% 0/0/1
= Detect MSTest v2 on VSTest and recommend migration path Eligible +0.0% +0.0% 0/1/0
= Execute targeted MSTest v2 to v3 migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect MSTest v2 on VSTest and recommend migration path: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.test-quality-auditor (claude-sonnet-5)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline request to generate new tests Excluded (activation contract) -100.0% -100.0% 0/0/1
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Route a curated test list to per-test decisions Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose test smells and propose a repair order: Both responses are substantively correct and actionable. A is slightly better because it stays tightly focused on the requested assessment and explains the regression-detection failures plainly. B is similarly thorough but is mildly diluted by an irrelevant test-execution cave...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assertion quality analysis Eligible +0.0% +0.0% 0/1/0
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
= Identify behavior gaps that existing tests would miss Eligible +0.0% +0.0% 0/1/0
= Route a curated test list to per-test decisions Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assertion quality analysis: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.testability-migration (claude-sonnet-5)

Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%

Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Migrate time dependencies and add deterministic tests Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Migrate time dependencies and add deterministic tests: Position-swap inconsistent (forward: skill, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.testability-migration (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan Eligible -100.0% -40.0% 0/0/1
= Targeted request: just migrate DateTime to TimeProvider Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: Response A delivers a cleaner, more idiomatic migration plan with better execution quality (0 errors vs 6). While Response B provides more precise line numbers, Response A's recommendations—using ILogger instead of IConsole, avoiding external dependencies, using the options pa...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — coverage-analysis (claude-sonnet-5)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +15.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 24 paired runs (12W/6T/6L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Preserve a classic packages.config project when coverage data is absent Eligible +0.0% +0.0% 1/0/1
▼ Reconcile a coverage target spread across several members Eligible -50.0% -20.0% 0/1/1
= Run coverage from scratch without existing data Eligible +0.0% -30.0% 1/0/1
▼ Stay dormant for one-member CRAP analysis Excluded (activation contract) -50.0% -20.0% 0/1/1
▼ Stay dormant for static source-to-test pairing Excluded (activation contract) -50.0% -20.0% 0/1/1
▼ Stay dormant for test trait distribution Excluded (activation contract) -50.0% -20.0% 0/1/1

Illustrative judge evidence:

  • Preserve a classic packages.config project when coverage data is absent: Both responses satisfy the constraints and reach the correct no-report conclusion. A is marginally stronger because it clearly explains the inability to calculate CRAP now and gives the complete legitimate next-step alternatives.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — coverage-analysis (gpt-5.6-luna)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +12.5% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded

Warnings: Activation: isolated 7/8; plugin 7/8

Overfit: Moderate (score 0.44)

Repeated-run reliability (not used by the gate): 24 paired runs (10W/10T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Distinguish partially covered branches from covered lines Eligible +50.0% +20.0% 1/1/0
= Project-wide coverage analysis with existing Cobertura data Eligible +0.0% +0.0% 0/2/0
▼ Reconcile a coverage target spread across several members Eligible -100.0% -40.0% 0/0/2
▼ Stay dormant for one-member CRAP analysis Excluded (activation contract) -100.0% -40.0% 0/0/2
= Stay dormant for static source-to-test pairing Excluded (activation contract) +0.0% +0.0% 0/2/0

Illustrative judge evidence:

  • Distinguish partially covered branches from covered lines: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — migrate-xunit-to-xunit-v3 (claude-sonnet-5)

Why: Net win +41.7% (7W/3T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +16.7% across 12 paired run(s) — not credible (sign test p=0.090 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=12; 7W/3T/2L; d=9; p=0.090; net +41.7%

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 12 paired runs (7W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Consolidate xunit.extensibility packages and remove xunit.abstractions Eligible +0.0% +0.0% 0/1/0
= Migrate project with YTest.MTP.XUnit2 to xUnit.net v3 preserving MTP Eligible +0.0% +0.0% 0/1/0
▼ Migrate xUnit v2 packages managed via Central Package Management Eligible -100.0% -40.0% 0/0/1
▼ Update Xunit.Combinatorial and Xunit.StaFact companion packages Eligible -100.0% -40.0% 0/0/1
= Update custom FactAttribute to include source information parameters Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Consolidate xunit.extensibility packages and remove xunit.abstractions: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — scaffold-dotnet-test-project (claude-sonnet-5)

Why: Net win +30.0% (5W/3T/2L over 10 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +18.0% across 10 paired run(s) — not credible (sign test p=0.227 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=10; 5W/3T/2L; d=7; p=0.227; net +30.0%

Warnings: Activation: isolated 9/10; plugin 8/10

Overfit: Low (score 0.19)

Repeated-run reliability (not used by the gate): 10 paired runs (5W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Add an existing test project to the CI solution filter Eligible -100.0% -40.0% 0/0/1
▼ Create the first pricing test project with central packages Eligible -100.0% -40.0% 0/0/1
▲ Include new tests in the solution filter used by CI Eligible +100.0% +40.0% 1/0/0
= Register an existing test project in the CI SDK solution Eligible +0.0% +0.0% 0/1/0
= Register the first tests in an SDK solution file Eligible +0.0% +0.0% 0/1/0
= Repair a missing production project reference Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add an existing test project to the CI solution filter: Both responses complete the requested minimal fix correctly and validate that the filter includes and executes the test project. A has a small edge because it explicitly verifies both the exact filter build and test commands; B's dotnet test validation is still substantively s...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-xunit-to-mstest (claude-sonnet-5)

Why: Net win +56.3% (11W/3T/2L over 16 preference-eligible stimulus vote(s), sign test p=0.011), mean preference +28.2% across 17 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=16; 11W/3T/2L; d=13; p=0.011; net +56.3%; 1 dormancy excluded

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 17 paired runs (11W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on xUnit v2 to v3 without framework conversion Excluded (activation contract) +0.0% +0.0% 0/1/0
= Convert IClassFixture to ClassInitialize Eligible +0.0% +0.0% 0/1/0
= Convert ITestOutputHelper to TestContext Eligible +0.0% +0.0% 0/1/0
▼ Handle ICollectionFixture explicitly (do not silently widen scope) Eligible -100.0% -40.0% 0/0/1
= Preserve type and sequence assertion semantics Eligible +0.0% +0.0% 0/1/0
▼ Preserve xUnit parallelization default with [assembly: Parallelize] Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Convert IClassFixture to ClassInitialize: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — scaffold-dotnet-test-project (gpt-5.6-luna)

Why: Net win +70.0% (7W/3T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +28.0% across 10 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=10; 7W/3T/0L; d=7; p=0.008; net +70.0%

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 10 paired runs (7W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add an existing test project to the CI solution filter Eligible +0.0% +0.0% 0/1/0
= Create the first pricing test project with central packages Eligible +0.0% +0.0% 0/1/0
= Register the first tests in an SDK solution file Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add an existing test project to the CI solution filter: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 4 results are in Full Results.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1289 in dotnet/skills, download eval artifacts with gh run download 37945406136 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/aa079b4c89962ca21fdbf471eda78784a1836590/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

16 model/target results across 8 targets and 2 models — ✅ 6 improved, ➖ 10 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 728658efd75a1270a1d844e42d43f96cf3dbcf2c; 2 judge models.

Measurement health: 16 expected / 16 observed / 16 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.test-engineer claude-sonnet-5 ➖ Improvement signal, unproven n=9; 6W/2T/1L; d=7; p=0.063; net +55.6% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-engineer gpt-5.6-luna ➖ Improvement signal, unproven n=9; 6W/2T/1L; d=7; p=0.063; net +55.6% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-migration claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-migration gpt-5.6-luna ➖ Improvement signal, unproven n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-quality-auditor claude-sonnet-5 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 1 dormancy excluded — — None.
agent.test-quality-auditor gpt-5.6-luna ➖ Improvement signal, unproven n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration claude-sonnet-5 ➖ Improvement signal, unproven n=5; 4W/0T/1L; d=5; p=0.188; net +60.0% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration gpt-5.6-luna ➖ Mixed evidence n=5; 1W/3T/1L; d=2; p=0.750; net +0.0% — — Compare winning and losing scenarios to isolate where the target helps versus hurts.
coverage-analysis claude-sonnet-5 ➖ Improvement signal, unproven n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded 🟡 0.43 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
coverage-analysis gpt-5.6-luna ✅ Improved n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%; 4 dormancy excluded 🟡 0.37 Activation: isolated 7/8; plugin 7/8 Fix activation gaps; Review overfit evidence.
migrate-xunit-to-mstest claude-sonnet-5 ✅ Improved n=16; 11W/5T/0L; d=11; p=0.000; net +68.8%; 1 dormancy excluded 🟡 0.31 — Review overfit evidence.
migrate-xunit-to-mstest gpt-5.6-luna ✅ Improved n=16; 11W/4T/1L; d=12; p=0.003; net +62.5%; 1 dormancy excluded — — None.
migrate-xunit-to-xunit-v3 claude-sonnet-5 ➖ Improvement signal, unproven n=12; 8W/2T/2L; d=10; p=0.055; net +50.0% 🟡 0.28 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-xunit-v3 gpt-5.6-luna ➖ Improvement signal, unproven n=12; 6W/3T/3L; d=9; p=0.254; net +25.0% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
scaffold-dotnet-test-project claude-sonnet-5 ✅ Improved n=10; 6W/4T/0L; d=6; p=0.016; net +60.0% ✅ 0.19 Activation: isolated 9/10; plugin 10/10 Fix activation gaps.
scaffold-dotnet-test-project gpt-5.6-luna ✅ Improved n=10; 7W/3T/0L; d=7; p=0.008; net +70.0% 🟡 0.43 — Review overfit evidence.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, unproven — agent.test-engineer (claude-sonnet-5)

Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +42.2% across 9 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%

Repeated-run reliability (not used by the gate): 9 paired runs (6W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate collaborating Go package tests Eligible +0.0% +0.0% 0/1/0
▼ Generate tests with validation execution unavailable Eligible -100.0% -40.0% 0/0/1
= Repair weak tests and close named gaps Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Generate collaborating Go package tests: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.test-engineer (gpt-5.6-luna)

Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +15.6% across 9 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%

Repeated-run reliability (not used by the gate): 9 paired runs (6W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate a project-wide pytest suite across modules Eligible +0.0% +0.0% 0/1/0
= Generate project-wide xUnit tests for a .NET library Eligible +0.0% +0.0% 0/1/0
▼ Generate tests with validation execution unavailable Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Generate a project-wide pytest suite across modules: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.test-migration (claude-sonnet-5)

Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +40.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 6 paired runs (3W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline request to write new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
= Detect MSTest v2 on VSTest and recommend migration path Eligible +0.0% +0.0% 0/1/0
= Plan a staged MSTest v2 to v4 and MTP migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect MSTest v2 on VSTest and recommend migration path: Position-swap inconsistent (forward: skill, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.test-migration (gpt-5.6-luna)

Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 6 paired runs (3W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline request to write new tests Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Detect xUnit v2 and route to xUnit v3 migration Eligible -100.0% -40.0% 0/0/1
▼ Execute targeted MSTest v2 to v3 migration Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Detect xUnit v2 and route to xUnit v3 migration: Both responses successfully completed the xUnit upgrade task with similar final outcomes (migrated to v4.0.1). However, Response A is slightly better because: (1) it explicitly mentioned updating Microsoft.NET.Test.Sdk to 18.10.1, which is a critical infrastructure upgrade for...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +42.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Identify behavior gaps that existing tests would miss Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose test smells and propose a repair order: Response A provides a more thorough, better-organized analysis with superior practical guidance. While both responses correctly identify the same core test smells and explain behavioral risks adequately, Response A excels in two key areas: (1) slightly more comprehensive probl...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.testability-migration (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +48.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%

Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Migrate time dependencies and add deterministic tests Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Migrate time dependencies and add deterministic tests: A completed nearly all requested implementation work, including the sibling xUnit project and deterministic boundary tests, while B completed only the production migration and left the required test project/tests uncreated. Both share the inability to execute tests, but A's de...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — agent.testability-migration (gpt-5.6-luna)

Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 5 paired run(s) — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%

Repeated-run reliability (not used by the gate): 5 paired runs (1W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Full pipeline: detect statics and recommend migration plan Eligible +0.0% +0.0% 0/1/0
▼ Inventory static dependencies without modifying the project Eligible -100.0% -40.0% 0/0/1
= Replace filesystem statics without touching unrelated dependencies Eligible +0.0% +0.0% 0/1/0
= Targeted request: just migrate DateTime to TimeProvider Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — coverage-analysis (claude-sonnet-5)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 24 paired runs (12W/9T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Preserve a classic packages.config project when coverage data is absent Eligible +0.0% +30.0% 1/0/1
= Project-wide coverage analysis with existing Cobertura data Eligible +0.0% +0.0% 0/2/0
▼ Refactoring safety assessment from coverage data Eligible -50.0% -20.0% 0/1/1
= Stay dormant for one-member CRAP analysis Excluded (activation contract) +0.0% +0.0% 1/0/1
= Stay dormant for static source-to-test pairing Excluded (activation contract) +0.0% +0.0% 0/2/0
= Stay dormant for test trait distribution Excluded (activation contract) +0.0% +0.0% 0/2/0

Illustrative judge evidence:

  • Preserve a classic packages.config project when coverage data is absent: Both responses are substantively compliant and reach the correct stop condition. A is marginally stronger because it directly closes the requested CRAP-risk portion by explaining why CRAP cannot be calculated without actual coverage data.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — migrate-xunit-to-xunit-v3 (claude-sonnet-5)

Why: Net win +50.0% (8W/2T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.055), mean preference +35.0% across 12 paired run(s) — not credible (sign test p=0.055 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=12; 8W/2T/2L; d=10; p=0.055; net +50.0%

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 12 paired runs (8W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Consolidate xunit.extensibility packages and remove xunit.abstractions Eligible -100.0% -40.0% 0/0/1
▼ Convert async void test methods to async Task Eligible -100.0% -40.0% 0/0/1
= Migrate xUnit v2 packages managed via Central Package Management Eligible +0.0% +0.0% 0/1/0
= Update BeforeAfterTestAttribute overrides with IXunitTest parameter Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Consolidate xunit.extensibility packages and remove xunit.abstractions: Both outcomes compile and run clean discovery, and both make the source-level migration. A better matches the task's explicit package-consolidation requirement by retaining a direct single xunit.v3.extensibility.core reference; B's MTP-off package choice is more indirect and l...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — migrate-xunit-to-xunit-v3 (gpt-5.6-luna)

Why: Net win +25.0% (6W/3T/3L over 12 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +20.0% across 12 paired run(s) — not credible (sign test p=0.254 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=12; 6W/3T/3L; d=9; p=0.254; net +25.0%

Repeated-run reliability (not used by the gate): 12 paired runs (6W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Consolidate xunit.extensibility packages and remove xunit.abstractions Eligible -100.0% -100.0% 0/0/1
= Convert string-based attribute constructors to typeof syntax Eligible +0.0% +0.0% 0/1/0
▼ Migrate Xunit.SkippableFact to xUnit.net v3 built-in skip APIs Eligible -100.0% -40.0% 0/0/1
= Recognize project already on xUnit.net v3 — no migration needed Eligible +0.0% +0.0% 0/1/0
▼ Update BeforeAfterTestAttribute overrides with IXunitTest parameter Eligible -100.0% -40.0% 0/0/1
= Update Xunit.Combinatorial and Xunit.StaFact companion packages Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Consolidate xunit.extensibility packages and remove xunit.abstractions: Response A directly fulfills the specific technical requirements stated in the rubric by using xunit.v3.extensibility.core as the consolidated package, whereas Response B takes an alternative approach with xunit.v3 that doesn't match the specific requirement. Response A also p...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — coverage-analysis (gpt-5.6-luna)

Why: Net win +62.5% (5W/3T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +19.2% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Fix activation gaps; Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%; 4 dormancy excluded

Warnings: Activation: isolated 7/8; plugin 7/8

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 24 paired runs (15W/4T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Distinguish partially covered branches from covered lines Eligible +0.0% +0.0% 1/0/1
= Project-wide coverage analysis with existing Cobertura data Eligible +0.0% +0.0% 0/2/0
= Reconcile a coverage target spread across several members Eligible +0.0% +0.0% 1/0/1
= Stay dormant for behavioral gap analysis Excluded (activation contract) +0.0% +0.0% 1/0/1
▼ Stay dormant for static source-to-test pairing Excluded (activation contract) -50.0% -20.0% 0/1/1
= Stay dormant for test trait distribution Excluded (activation contract) +0.0% +0.0% 1/0/1

Illustrative judge evidence:

  • Distinguish partially covered branches from covered lines: Response A provides a slightly more rigorous and careful answer. It uses conditional language ('where applicable') to avoid overstating what must be tested, explicitly names short-circuit operands as a consideration, and maintains better abstraction throughout without inventin...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — migrate-xunit-to-mstest (claude-sonnet-5)

Why: Net win +68.8% (11W/5T/0L over 16 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +38.8% across 17 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=16; 11W/5T/0L; d=11; p=0.000; net +68.8%; 1 dormancy excluded

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 17 paired runs (12W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Convert IClassFixture to ClassInitialize Eligible +0.0% +0.0% 0/1/0
= Convert ITestOutputHelper to TestContext Eligible +0.0% +0.0% 0/1/0
= Convert MemberData and TheoryData to DynamicData Eligible +0.0% +0.0% 0/1/0
= Preserve effective assembly and class owners on methods Eligible +0.0% +0.0% 0/1/0
= Preserve type and sequence assertion semantics Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Convert IClassFixture to ClassInitialize: Both perform a correct, verified xUnit-to-MSTest conversion and satisfy every specified fixture-lifecycle requirement. Their package choices differ but are both valid and tested successfully.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — scaffold-dotnet-test-project (claude-sonnet-5)

Why: Net win +60.0% (6W/4T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +36.0% across 10 paired run(s) — credibly better

Next action: Fix activation gaps.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=10; 6W/4T/0L; d=6; p=0.016; net +60.0%

Warnings: Activation: isolated 9/10; plugin 10/10

Overfit: Low (score 0.19)

Repeated-run reliability (not used by the gate): 10 paired runs (6W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add an existing test project to the CI solution filter Eligible +0.0% +0.0% 0/1/0
= Register an existing test project omitted from the solution Eligible +0.0% +0.0% 0/1/0
= Register the first tests in an SDK solution file Eligible +0.0% +0.0% 0/1/0
= Repair a missing production project reference Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add an existing test project to the CI solution filter: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — scaffold-dotnet-test-project (gpt-5.6-luna)

Why: Net win +70.0% (7W/3T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +28.0% across 10 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=10; 7W/3T/0L; d=7; p=0.008; net +70.0%

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 10 paired runs (7W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add an existing test project to the CI solution filter Eligible +0.0% +0.0% 0/1/0
= Keep a domain-only request bounded Eligible +0.0% +0.0% 0/1/0
= Reuse an existing suitable test project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add an existing test project to the CI solution filter: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 2 results are in Full Results.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1289 in dotnet/skills, download eval artifacts with gh run download 37952260334 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/728658efd75a1270a1d844e42d43f96cf3dbcf2c/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 9, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 728658e. cc @dotnet/dotnet-testing — please review.

@Evangelink
Evangelink merged commit 0478872 into main Oct 9, 2026
47 checks passed
@Evangelink
Evangelink deleted the dev/amauryleve/portable-review-corrections branch October 9, 2026 19:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-review PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants