You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Fix portable test guidance review regressions - #1289
Contribute portable TestFx review corrections back to the source repository, after comparing current upstream:
Retain explicitly requested Python regressions and required assertions despite production/native-import/tooling blockers; remove only speculative out-of-scope cases.
Select .NET discovery from actual runner/command mode, distinguishing positional VSTest, positional bridged MTP and documented named native MTP. Require successful discovery exit and runner-specific output, not masked counting pipelines.
Recognize .slnx; preserve explicit .sln/.slnx/.slnf/.csproj; stop on ambiguity/lookup errors; bound Git inventory to requested/containing roots; resolve the exact selected collection graph before classification.
Preserve xUnit executable/native runner configuration, reserve TestingPlatformDotnetTestSupport for the VSTest-command-mode bridge, and use mode-aware direct-project/solution verification, retaining classic native runners.
Preserve assertion packages/imports and correct AwesomeAssertions namespace. Map effective class/assembly Owner traits to method-only MSTest attributes, deduplicate and flag conflicts/shared inheritance.
Replace implicit migration commits with phase validation and explicit-user-only commits.
Classify xUnit v3 reader rejection. Alternate permitted readers require positively confirmed technical/path-normalization limitations plus authorized access. Stop/report permission/policy/content-exclusion denials without tool/shell/alias/agent bypass or reconstruction. Unknown rejection is unavailable, not missing; disclose incomplete work.
Add owning migration fixtures/goldens/authenticated graders for Owner, assertion-library/import/original-assertion preservation and evaluated VSTest state. Remove conflict-prompt answer cues; add an uncued xUnit version-upgrade dormancy guard. Cover complete read-only inventories and all original test assertions, not merely test names/attributes or .Should() presence.
Preserve Ruby harness runner failures: use same-process fail-fast Bash, RSpec dry-run summaries or the configured real Minitest task and test summary, never counting pipelines, suppressed stderr or post-failure runner fallback.
Scope/rollout: Helper-loading/routing, including filter-syntax, overlaps pending #1287 and is untouched. Open #1189 concerns legacy generation tracking/project placement, not these defects. Current test-engineer already has its scratch contract; no deleted generator is restored. TestFx-specific legacy-generator/BOM policies are not portable and remain unchanged upstream. TestFx still pins e468462d8a0900f5278331c1ae12f45e015bb656: this contribution is not yet consumed by that pin. Merged-source refresh and later removal of applicable exceptions remain downstream work. TestFx-specific adaptations to removed vendored migration plugins are not propagated; valid source framework/platform routes remain unchanged.
Final exact files: 16/16 discovery tests passed; 17/17 portable guidance tests passed, no skips; actual Pester5.7.1 ran9/9. The parent's original12+7 run and earlier14-test portable run are historical, not the current submitted portable count. Coverage includes entry/ambiguity/lookup/classic classification/root isolation/exact graph; promised invocation/error/non-persistence agreement; three reader rejection dispositions; and three Ruby harness regressions. All three reader guards reject upstream's unconditional fallback plus targeted unsafe mutations (6/6 rejected). These are guidance-contract checks, not simulated access-control enforcement; no denied file was accessed. Subprocesses are timeout-bounded; absent Pester-v5 is explicitly skipped/not executed.
Ruby harness follow-up (728658efd): The portable unittest command above was rerun with a 180s outer subprocess bound: exit 0, 17/17, no skips, including actual Pester9/9. Three new tests check selection/count guidance and execute the two published Bash blocks with synthetic runners: four executions, expected exits 0,0,23,23, preserved stdout summaries/stderr, and exactly one selected runner per invocation (30s bound each). GNU Bash5.3.15(2)-release on Git for Windows. Both unchanged upstream commands were separately reproduced returning Bash exit 0 despite runner exit 23. Eval-quality against origin/main and whitespace check exit 0, existing warnings only. No eval YAML or framework/plugin routes changed in this follow-up. Ruby/Bundler are absent: synthetic exit propagation is not native RSpec/Rake/Minitest integration proof; no tools installed.
Ruby/RSpec and Kotlin/Gradle/JVM tools are absent. rustc --version and cargo --version each exit 2, tool_not_installed; installed launchers are not a usable toolchain. Those three samples have structural evidence only. No follow-up tooling/global settings were installed or changed.
All exit 0: plugin checks 23 skills/10 agents, 6 skills/1 agent, eval-quality with existing warnings only, whitespace clean. Both plugins/regression files were rerun after the reader correction; eval-quality/whitespace repeated after final Owner grading. Validator restore followed a proven missing-assets failure. Owning eval replaces deprecated config: with defaults: and adds capability/risk/journey tags; rubric weights, plugin metadata and external domains are unchanged.
Native/reference contracts: Original xUnit fixtures ran 3/3,3/3,2/2 tests. Migrated goldens ran 3/3,2/2 on MSTest4.4.1, Test.Sdk18.9.0, SDK11.0.100-rc.1.26425.128, net8.0/VSTest. Inputs: xunit.v3 3.2.2, adapter 3.1.5, Test.Sdk 17.13.0. AwesomeAssertions 9.6.0, local/global/project imports and original assertions survive. Owner reflection requires one alice Owner on each of exactly three discoverable methods. Helper digests and complete staged input hashes are authenticated.
Correct states exit 0, with exact nonzero test counts/markers. Windows replay resolves specification python3 to installed python. Final replay accepts all seven deterministic grader invocations against three reference states. Initial seven mutations reject lost/wrong/duplicate Owner, silent conflict rewriting and separate local/global/project FluentAssertions imports. Runner replay: 17 commands, package VSTest and MSTest.Sdk4.4.1 + UseVSTest=true accepted, SDK references run 3/2 tests, six native/bridged-MTP mutations and a boundary source edit rejected.
Final assertion/inventory replay: 18 commands, arrow/block formatting accepted, four weakened-assertion/comment-decoy mutations rejected, both complete read-only goldens accepted, six add-file mutations rejected (Directory.Build.props, Directory.Build.targets and arbitrary JSON in each case). Inventory covers every file outside known .eval/.git metadata and bin/obj, with exact supplied hashes, not extension filtering.
Final Owner assertion guard:verify_owner_assertions.py exits 0 across nine native command logs. Original three assertion bodies are required alongside reflected names/Owner/discovery metadata. Golden and three individual block-format variants pass. Removing each assertion independently or emptying all three bodies produces four rejected mutations (exit1). The final seven-grader replay, production parser and full-PR eval-quality pass after the authenticated Owner digest update. Logs/scripts/reference fixtures are retained in migration-review-verified, migration-review-final-graders, runner-review, final-eval-guards, owner-assertion-review and final quality logs.
Production probe — executed, non-passing: Installed VallyCLI0.9.0, not claimed equivalent to documented CI0.14:
vally lint plugins\dotnet-test-migration\skills\migrate-xunit-to-mstest -e tests\dotnet-test-migration\migrate-xunit-to-mstest\eval.yaml
vally experiment run <retained-session-artifacts>\production-review\experiment.yaml --dry-run --workers 8
vally experiment run <retained-session-artifacts>\production-review\experiment.yaml --workers 8--output-dir <retained-session-artifacts>\production-review\results
Lint/dry run exit 0 (one default-scoring warning). Actual run exits 1 after 8/8 executor records completed successfully, zero executor errors/timeouts: four stimuli × baseline/isolated × one run, claude-opus-4.6 executor/judge, normal 8 workers, declared 20m timeout, outer1320s bound. Runtime used f5af583de; subsequent 82da18db8/aa079b4c8 grader strengthening was re-linted and deterministically replayed, not model-rerun.
Authenticated command graders returned 9009 because Windows python3 resolves to an uninstalled Microsoft Store alias. That host prerequisite was classified before content changes. All three positive isolated trials invoked the target; advisory trial did not, but its rubric caught an unwanted MSTest recommendation. The judge also flagged a possible project-import omission, not certified by the blocked deterministic grader. Neither arm is green; no preference W/T/L/powered verdict is claimed. No plugin arm, second family, rerun or broad matrix. Owning suite: 16 preference stimuli + one dormant boundary; four-case probe is deliberately below power floor. Raw results/trajectories/metadata/run.log retained in production-review\results\2026-10-09T13-12-26-559Z, run ID 8ef8bd30-2836-4785-a8cf-b482aa8770b7; Vally cleaned temporary model workspaces.
Real selected collection: Selected three-test project references another project with an intentionally failing test. dotnet test TestProject.csproj --logger "console;verbosity=normal" and dotnet test Selected.slnf --logger "console;verbosity=normal" both exit 0, run only the selected three tests, never the failing dependency. Published inventory agrees. Retained verify_collection_graph.py/collection-graph contain fixture/logs.
Composed runner discovery/execution: Every command exits 0; discovery contains real Invoice_is_discovered; direct/solution runs execute it.
Mode/components
Actual fixture commands
VSTest: SDK10.0.401, MSTest3.8.0, Test.Sdk17.13.0
dotnet test Invoice.Tests.csproj --no-build; dotnet test Billing.sln --list-tests --no-build; dotnet test Billing.sln --no-build
Bridged MTP: SDK8.0.425, xunit.v3 3.2.2, MTP1.9.1
dotnet test Invoice.Tests.csproj --no-build; dotnet test Billing.sln --no-build -p:TestingPlatformCaptureOutput=false -- --list-tests; dotnet test Billing.sln --no-build
Native MTP: SDK10.0.401, xunit.v3 3.2.2, MTP1.9.1, no bridge property
dotnet test --project Invoice.Tests.csproj --no-build; dotnet test --solution Billing.slnx --list-tests --no-build; dotnet test --solution Billing.slnx --no-build
Qualification: SDK10.0.401 also accepted positional native project syntax (dotnet test Invoice.Tests.csproj --no-build). Guidance aligns documented named forms, not universal parser failure; the initial negative fixture expectation was corrected.
Namespace/attribute probes: AwesomeAssertions9.6.0 with MSTest3.8.0/exact4.4.1. dotnet run -v:q and dotnet run --project <retained-session-artifacts>\native-contracts\metadata\Metadata.csproj -v:q exit 0; invalid class Owner via dotnet build --no-restore -v:q fails CS0592.
C++ native sample: Published blocks/minimal interface header compiled as C++20, MSVC19.51.36260.0, Catch2v3.8.1 commit 2b60af89e23d28eefc081bc930831ee9d45ea58b, after installed vcvars64.bat setup:
All exit 0: 2 cases/9 sections/12 assertions, CTest1/1. NMake ignores parallel; incidental missing-vswhere warning did not prevent build/execution. Missing Catch2 was established before obtaining the pinned dependency. Fixture/log/XML evidence retained.
Original Pester enum failure, .slnx misselection and sibling leakage were reproduced before correction. These are local native/reference checks, not CI results or certified model-quality improvement. No broad Vally/model-family matrix was launched.
Checklist
I searched existing issues and pull requests to avoid duplicates.
I kept this pull request focused and avoided unrelated refactors.
I added or updated tests, evals, or documentation when changing skill or agent behavior.
I updated CODEOWNERS when adding or moving owned content. (N/A: existing ownership of both affected plugins/tests covers all paths.)
I updated all marketplace manifests when plugin metadata changed. (N/A: no plugin metadata changed.)
I updated eng/known-domains.txt for any new external domains referenced by skill content. (N/A: no new domains.)
Evaluation changes only
For every eval-related change, including a new eval:
Three correctness scenarios and one uncued dormancy guard; existing already-MSTest no-op retained. Suite16 preference + one dormancy; four-case probe is not powered quality evidence.
Scenarios are necessary, distinct, and use natural prompts.
Graders cover the full result, accept the golden result, and reject a realistic mutation.
No-op, dormancy, and statistical power are covered where needed. (No-op/dormancy covered; powered quality not measured by this subset.)
I ran the applicable production evaluation path and recorded the result above. (Subset completed; exit1, revision/host/quality limitations reported, not a passing verdict.)
If this fixes a failed eval, I classified the failure before editing skill content. (Windows prerequisite classified; no prose changed in response to that run.)
If this broadly changes routing or behavior, I checked separate model-family evidence. (N/A: no broad routing change or matrix.)
👋 @Evangelink — this PR has 3 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
The bridged-MTP command omits -p:TestingPlatformCaptureOutput=false, so the bridge's default capture can hide the runner-owned test list that this section requires inspecting. The composed validation in the PR description used that property, so it did not validate the command published here. Include the property (or specify and verify the captured discovery artifact) and update the regression assertion to exercise the exact documented command.
This distinguishes “not in a repository” from a real Git failure by matching Git's English stderr. Git localizes this diagnostic, so the same standalone directory is treated as a fatal lookup error on non-English hosts, contrary to the portable no-repository path below. Force Git's locale to C while capturing this command (restoring the caller's environment afterward), or classify absence without parsing localized text.
Restrict reader fallback to confirmed authorized technical limitations,
stop on access denials, and disclose unavailable paths and incomplete work.
Guard the three rejection dispositions with focused guidance regressions.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
Both test-project predicates omit established markers such as MSTest.Sdk, Microsoft.Testing.Platform, and <IsTestProject>true</IsTestProject> (compare plugins/dotnet-test/skills/find-untested-sources/SKILL.md:143-148). A neutral-name project using one of those forms is removed from a selected solution graph and can incorrectly produce TEST_PROJECTS:0. Include these markers in both predicates and add a neutral-name MSTest.Sdk regression.
Match both original assertion bodies with formatting flexibility and
inventory every non-evaluator project file in read-only migration cases.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
When $root is a directory, $entry can be the only discovered .csproj even if that project is production code. This branch then limits $projectPaths to that non-test project and never performs the documented requested-directory/containing-Git-root test inventory. For example, requesting src/App when it contains only App.csproj and the repository keeps App.Tests.csproj under tests/ now reports TEST_PROJECTS:0 and stops. Distinguish an explicitly selected test entry from an automatically discovered non-test project, and add this layout to the executable regression.
These output graders do not prove the owner mappings required by the rubric. A response such as “alice, bob, and charlie conflict; Inherited is shared; manual review is required” passes all four patterns while omitting DirectTests.ConflictingOwners and failing to associate bob/charlie with BillingTests.Inherited/PaymentTests.Inherited. Add deterministic checks tying each conflict to its affected method so this realistic incomplete assessment is rejected.
The dormancy contract deterministically checks only that the answer mentions xUnit v3. A response can recommend converting to MSTest as well, pass this output grader and the unchanged-file check, yet violate the routing invariant; the reported bounded run actually exhibited this failure and only the LLM rubric caught it. Add a deterministic output contract that rejects an MSTest conversion recommendation while allowing a concise contrast such as “not MSTest.”
Require all three original assertion bodies alongside reflected method
and Owner metadata, with expression/block formatting preserved.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0b59cd01-f12f-4669-afb6-74b314e8b79a
This deterministic grader is order- and line-break-sensitive: a correct response such as Owner conflicts:\n- alice and bob ... fails because owner appears before conflict and . cannot cross the newline. Accept either ordering across lines so valid conflict reports are not rejected.
This rubric scores whether the target skill was invoked, which is harness metadata rather than an answer outcome and biases the pairwise judge toward skill-specific vocabulary. expect_activation: false already enforces dormancy; keep the rubric limited to the xUnit-v3 answer, absence of MSTest conversion, and unchanged files.
Published bridged command omits output capture disabling property
The bridged-MTP command does not expose the test list that the following paragraph requires reviewers to inspect: dotnet test captures the MTP application output unless TestingPlatformCaptureOutput is disabled. The PR's own validated bridged command includes -p:TestingPlatformCaptureOutput=false, but this published command omits it, so it can exit successfully without giving the identities/count needed for the discovery check. Add that property here (or document the runner-owned artifact to inspect) and update the regression assertion accordingly.
This treats only Git's English “not a git repository” diagnostic as the normal no-repository case. Git diagnostics are localized, so the same lookup on a non-English host falls through to Git root lookup failed and aborts discovery instead of continuing with the requested directory. Classify the no-repository state without matching localized stderr (for example, by checking for a containing .git marker while still surfacing broken-marker/lookup failures).
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.test-engineer
claude-sonnet-5
✅ Improved
n=9; 7W/1T/1L; d=8; p=0.035; net +66.7%
—
—
None.
agent.test-engineer
gpt-5.6-luna
✅ Improved
n=9; 7W/1T/1L; d=8; p=0.035; net +66.7%
—
—
None.
agent.test-migration
claude-sonnet-5
➖ Improvement signal, unproven
n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-migration
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-quality-auditor
claude-sonnet-5
➖ Improvement signal, unproven
n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-quality-auditor
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
coverage-analysis
claude-sonnet-5
➖ Improvement signal, unproven
n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded
🟡 0.24
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
coverage-analysis
gpt-5.6-luna
➖ Improvement signal, unproven
n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded
🟡 0.44
Activation: isolated 7/8; plugin 7/8
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-mstest
claude-sonnet-5
✅ Improved
n=16; 11W/3T/2L; d=13; p=0.011; net +56.3%; 1 dormancy excluded
🟡 0.41
—
Review overfit evidence.
migrate-xunit-to-mstest
gpt-5.6-luna
✅ Improved
n=16; 13W/2T/1L; d=14; p=0.001; net +75.0%; 1 dormancy excluded
—
—
None.
migrate-xunit-to-xunit-v3
claude-sonnet-5
➖ Improvement signal, unproven
n=12; 7W/3T/2L; d=9; p=0.090; net +41.7%
🟡 0.28
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-xunit-v3
gpt-5.6-luna
✅ Improved
n=12; 11W/1T/0L; d=11; p=0.000; net +91.7%
—
—
None.
scaffold-dotnet-test-project
claude-sonnet-5
➖ Improvement signal, unproven
n=10; 5W/3T/2L; d=7; p=0.227; net +30.0%
✅ 0.19
Activation: isolated 9/10; plugin 8/10
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
scaffold-dotnet-test-project
gpt-5.6-luna
✅ Improved
n=10; 7W/3T/0L; d=7; p=0.008; net +70.0%
🟡 0.48
—
Review overfit evidence.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +66.7% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 6 paired runs (5W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Plan a staged MSTest v2 to v4 and MTP migration
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Plan a staged MSTest v2 to v4 and MTP migration: Both are directionally correct and satisfy the core requested staging. A is more actionable and safety-oriented, particularly through its explicit baseline stage, regression comparisons, and more detailed gates. Its proposed MTP package wording is somewhat less precise than id...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +33.3% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline request to generate new tests
Excluded (activation contract)
-100.0%
-100.0%
0/0/1
▼ Diagnose test smells and propose a repair order
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a curated test list to per-test decisions
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose test smells and propose a repair order: Both responses are substantively correct and actionable. A is slightly better because it stays tightly focused on the requested assessment and explains the regression-detection failures plainly. B is similarly thorough but is mildly diluted by an irrelevant test-execution cave...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan
Eligible
-100.0%
-40.0%
0/0/1
= Targeted request: just migrate DateTime to TimeProvider
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Full pipeline: detect statics and recommend migration plan: Response A delivers a cleaner, more idiomatic migration plan with better execution quality (0 errors vs 6). While Response B provides more precise line numbers, Response A's recommendations—using ILogger instead of IConsole, avoiding external dependencies, using the options pa...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +15.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 24 paired runs (12W/6T/6L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Preserve a classic packages.config project when coverage data is absent
Eligible
+0.0%
+0.0%
1/0/1
▼ Reconcile a coverage target spread across several members
Eligible
-50.0%
-20.0%
0/1/1
= Run coverage from scratch without existing data
Eligible
+0.0%
-30.0%
1/0/1
▼ Stay dormant for one-member CRAP analysis
Excluded (activation contract)
-50.0%
-20.0%
0/1/1
▼ Stay dormant for static source-to-test pairing
Excluded (activation contract)
-50.0%
-20.0%
0/1/1
▼ Stay dormant for test trait distribution
Excluded (activation contract)
-50.0%
-20.0%
0/1/1
Illustrative judge evidence:
Preserve a classic packages.config project when coverage data is absent: Both responses satisfy the constraints and reach the correct no-report conclusion. A is marginally stronger because it clearly explains the inability to calculate CRAP now and gives the complete legitimate next-step alternatives.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +12.5% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +41.7% (7W/3T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +16.7% across 12 paired run(s) — not credible (sign test p=0.090 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +30.0% (5W/3T/2L over 10 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +18.0% across 10 paired run(s) — not credible (sign test p=0.227 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=10; 5W/3T/2L; d=7; p=0.227; net +30.0%
Warnings: Activation: isolated 9/10; plugin 8/10
Overfit: Low (score 0.19)
Repeated-run reliability (not used by the gate): 10 paired runs (5W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Add an existing test project to the CI solution filter
Eligible
-100.0%
-40.0%
0/0/1
▼ Create the first pricing test project with central packages
Eligible
-100.0%
-40.0%
0/0/1
▲ Include new tests in the solution filter used by CI
Eligible
+100.0%
+40.0%
1/0/0
= Register an existing test project in the CI SDK solution
Eligible
+0.0%
+0.0%
0/1/0
= Register the first tests in an SDK solution file
Eligible
+0.0%
+0.0%
0/1/0
= Repair a missing production project reference
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add an existing test project to the CI solution filter: Both responses complete the requested minimal fix correctly and validate that the filter includes and executes the test project. A has a small edge because it explicitly verifies both the exact filter build and test commands; B's dotnet test validation is still substantively s...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +56.3% (11W/3T/2L over 16 preference-eligible stimulus vote(s), sign test p=0.011), mean preference +28.2% across 17 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Why: Net win +70.0% (7W/3T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +28.0% across 10 paired run(s) — credibly better
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1289 in dotnet/skills, download eval artifacts with gh run download 37945406136 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/aa079b4c89962ca21fdbf471eda78784a1836590/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.test-engineer
claude-sonnet-5
➖ Improvement signal, unproven
n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-engineer
gpt-5.6-luna
➖ Improvement signal, unproven
n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-migration
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%; 1 dormancy excluded
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-migration
gpt-5.6-luna
➖ Improvement signal, unproven
n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.test-quality-auditor
claude-sonnet-5
✅ Improved
n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 1 dormancy excluded
—
—
None.
agent.test-quality-auditor
gpt-5.6-luna
➖ Improvement signal, unproven
n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration
claude-sonnet-5
➖ Improvement signal, unproven
n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration
gpt-5.6-luna
➖ Mixed evidence
n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%
—
—
Compare winning and losing scenarios to isolate where the target helps versus hurts.
coverage-analysis
claude-sonnet-5
➖ Improvement signal, unproven
n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded
🟡 0.43
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
coverage-analysis
gpt-5.6-luna
✅ Improved
n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%; 4 dormancy excluded
🟡 0.37
Activation: isolated 7/8; plugin 7/8
Fix activation gaps; Review overfit evidence.
migrate-xunit-to-mstest
claude-sonnet-5
✅ Improved
n=16; 11W/5T/0L; d=11; p=0.000; net +68.8%; 1 dormancy excluded
🟡 0.31
—
Review overfit evidence.
migrate-xunit-to-mstest
gpt-5.6-luna
✅ Improved
n=16; 11W/4T/1L; d=12; p=0.003; net +62.5%; 1 dormancy excluded
—
—
None.
migrate-xunit-to-xunit-v3
claude-sonnet-5
➖ Improvement signal, unproven
n=12; 8W/2T/2L; d=10; p=0.055; net +50.0%
🟡 0.28
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
migrate-xunit-to-xunit-v3
gpt-5.6-luna
➖ Improvement signal, unproven
n=12; 6W/3T/3L; d=9; p=0.254; net +25.0%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
scaffold-dotnet-test-project
claude-sonnet-5
✅ Improved
n=10; 6W/4T/0L; d=6; p=0.016; net +60.0%
✅ 0.19
Activation: isolated 9/10; plugin 10/10
Fix activation gaps.
scaffold-dotnet-test-project
gpt-5.6-luna
✅ Improved
n=10; 7W/3T/0L; d=7; p=0.008; net +70.0%
🟡 0.43
—
Review overfit evidence.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +42.2% across 9 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +15.6% across 9 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +40.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 6 paired runs (3W/0T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline request to write new tests
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Detect xUnit v2 and route to xUnit v3 migration
Eligible
-100.0%
-40.0%
0/0/1
▼ Execute targeted MSTest v2 to v3 migration
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Detect xUnit v2 and route to xUnit v3 migration: Both responses successfully completed the xUnit upgrade task with similar final outcomes (migrated to v4.0.1). However, Response A is slightly better because: (1) it explicitly mentioned updating Microsoft.NET.Test.Sdk to 18.10.1, which is a critical infrastructure upgrade for...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +42.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Diagnose test smells and propose a repair order
Eligible
-100.0%
-40.0%
0/0/1
▼ Identify behavior gaps that existing tests would miss
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose test smells and propose a repair order: Response A provides a more thorough, better-organized analysis with superior practical guidance. While both responses correctly identify the same core test smells and explain behavioral risks adequately, Response A excels in two key areas: (1) slightly more comprehensive probl...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +48.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Migrate time dependencies and add deterministic tests
Eligible
-100.0%
-100.0%
0/0/1
Illustrative judge evidence:
Migrate time dependencies and add deterministic tests: A completed nearly all requested implementation work, including the sibling xUnit project and deterministic boundary tests, while B completed only the production migration and left the required test project/tests uncreated. Both share the inability to execute tests, but A's de...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 5 paired run(s) — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 24 paired runs (12W/9T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Preserve a classic packages.config project when coverage data is absent
Eligible
+0.0%
+30.0%
1/0/1
= Project-wide coverage analysis with existing Cobertura data
Eligible
+0.0%
+0.0%
0/2/0
▼ Refactoring safety assessment from coverage data
Eligible
-50.0%
-20.0%
0/1/1
= Stay dormant for one-member CRAP analysis
Excluded (activation contract)
+0.0%
+0.0%
1/0/1
= Stay dormant for static source-to-test pairing
Excluded (activation contract)
+0.0%
+0.0%
0/2/0
= Stay dormant for test trait distribution
Excluded (activation contract)
+0.0%
+0.0%
0/2/0
Illustrative judge evidence:
Preserve a classic packages.config project when coverage data is absent: Both responses are substantively compliant and reach the correct stop condition. A is marginally stronger because it directly closes the requested CRAP-risk portion by explaining why CRAP cannot be calculated without actual coverage data.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +50.0% (8W/2T/2L over 12 preference-eligible stimulus vote(s), sign test p=0.055), mean preference +35.0% across 12 paired run(s) — not credible (sign test p=0.055 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=12; 8W/2T/2L; d=10; p=0.055; net +50.0%
Overfit: Moderate (score 0.28)
Repeated-run reliability (not used by the gate): 12 paired runs (8W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Consolidate xunit.extensibility packages and remove xunit.abstractions
Eligible
-100.0%
-40.0%
0/0/1
▼ Convert async void test methods to async Task
Eligible
-100.0%
-40.0%
0/0/1
= Migrate xUnit v2 packages managed via Central Package Management
Eligible
+0.0%
+0.0%
0/1/0
= Update BeforeAfterTestAttribute overrides with IXunitTest parameter
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Consolidate xunit.extensibility packages and remove xunit.abstractions: Both outcomes compile and run clean discovery, and both make the source-level migration. A better matches the task's explicit package-consolidation requirement by retaining a direct single xunit.v3.extensibility.core reference; B's MTP-off package choice is more indirect and l...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +25.0% (6W/3T/3L over 12 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +20.0% across 12 paired run(s) — not credible (sign test p=0.254 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=12; 6W/3T/3L; d=9; p=0.254; net +25.0%
Repeated-run reliability (not used by the gate): 12 paired runs (6W/3T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Consolidate xunit.extensibility packages and remove xunit.abstractions
Eligible
-100.0%
-100.0%
0/0/1
= Convert string-based attribute constructors to typeof syntax
Eligible
+0.0%
+0.0%
0/1/0
▼ Migrate Xunit.SkippableFact to xUnit.net v3 built-in skip APIs
Eligible
-100.0%
-40.0%
0/0/1
= Recognize project already on xUnit.net v3 — no migration needed
Eligible
+0.0%
+0.0%
0/1/0
▼ Update BeforeAfterTestAttribute overrides with IXunitTest parameter
Eligible
-100.0%
-40.0%
0/0/1
= Update Xunit.Combinatorial and Xunit.StaFact companion packages
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Consolidate xunit.extensibility packages and remove xunit.abstractions: Response A directly fulfills the specific technical requirements stated in the rubric by using xunit.v3.extensibility.core as the consolidated package, whereas Response B takes an alternative approach with xunit.v3 that doesn't match the specific requirement. Response A also p...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — coverage-analysis (gpt-5.6-luna)
Why: Net win +62.5% (5W/3T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +19.2% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Fix activation gaps; Review overfit evidence.
Repeated-run reliability (not used by the gate): 24 paired runs (15W/4T/5L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Distinguish partially covered branches from covered lines
Eligible
+0.0%
+0.0%
1/0/1
= Project-wide coverage analysis with existing Cobertura data
Eligible
+0.0%
+0.0%
0/2/0
= Reconcile a coverage target spread across several members
Eligible
+0.0%
+0.0%
1/0/1
= Stay dormant for behavioral gap analysis
Excluded (activation contract)
+0.0%
+0.0%
1/0/1
▼ Stay dormant for static source-to-test pairing
Excluded (activation contract)
-50.0%
-20.0%
0/1/1
= Stay dormant for test trait distribution
Excluded (activation contract)
+0.0%
+0.0%
1/0/1
Illustrative judge evidence:
Distinguish partially covered branches from covered lines: Response A provides a slightly more rigorous and careful answer. It uses conditional language ('where applicable') to avoid overstating what must be tested, explicitly names short-circuit operands as a consideration, and maintains better abstraction throughout without inventin...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +68.8% (11W/5T/0L over 16 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +38.8% across 17 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Repeated-run reliability (not used by the gate): 17 paired runs (12W/5T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Convert IClassFixture to ClassInitialize
Eligible
+0.0%
+0.0%
0/1/0
= Convert ITestOutputHelper to TestContext
Eligible
+0.0%
+0.0%
0/1/0
= Convert MemberData and TheoryData to DynamicData
Eligible
+0.0%
+0.0%
0/1/0
= Preserve effective assembly and class owners on methods
Eligible
+0.0%
+0.0%
0/1/0
= Preserve type and sequence assertion semantics
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Convert IClassFixture to ClassInitialize: Both perform a correct, verified xUnit-to-MSTest conversion and satisfy every specified fixture-lifecycle requirement. Their package choices differ but are both valid and tested successfully.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +60.0% (6W/4T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +36.0% across 10 paired run(s) — credibly better
Why: Net win +70.0% (7W/3T/0L over 10 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +28.0% across 10 paired run(s) — credibly better
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1289 in dotnet/skills, download eval artifacts with gh run download 37952260334 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/728658efd75a1270a1d844e42d43f96cf3dbcf2c/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Contribute portable TestFx review corrections back to the source repository, after comparing current upstream:
.slnx; preserve explicit.sln/.slnx/.slnf/.csproj; stop on ambiguity/lookup errors; bound Git inventory to requested/containing roots; resolve the exact selected collection graph before classification.TestingPlatformDotnetTestSupportfor the VSTest-command-mode bridge, and use mode-aware direct-project/solution verification, retaining classic native runners..Should()presence.Commits: original
f3d02293e; first review5d7609327; examplesacf7fb308; runner/routingf5af583de; reader policyeccaebf25; assertion/inventory82da18db8; Owner assertionsaa079b4c89962ca21fdbf471eda78784a1836590; Ruby harness728658efd75a1270a1d844e42d43f96cf3dbcf2c.Scope/rollout: Helper-loading/routing, including
filter-syntax, overlaps pending #1287 and is untouched. Open #1189 concerns legacy generation tracking/project placement, not these defects. Currenttest-engineeralready has its scratch contract; no deleted generator is restored. TestFx-specific legacy-generator/BOM policies are not portable and remain unchanged upstream. TestFx still pinse468462d8a0900f5278331c1ae12f45e015bb656: this contribution is not yet consumed by that pin. Merged-source refresh and later removal of applicable exceptions remain downstream work. TestFx-specific adaptations to removed vendored migration plugins are not propagated; valid source framework/platform routes remain unchanged.Related issue
N/A — found during review of microsoft/testfx#11856. No issue closure is implied.
Validation
Final exact files: 16/16 discovery tests passed; 17/17 portable guidance tests passed, no skips; actual Pester5.7.1 ran9/9. The parent's original12+7 run and earlier14-test portable run are historical, not the current submitted portable count. Coverage includes entry/ambiguity/lookup/classic classification/root isolation/exact graph; promised invocation/error/non-persistence agreement; three reader rejection dispositions; and three Ruby harness regressions. All three reader guards reject upstream's unconditional fallback plus targeted unsafe mutations (6/6 rejected). These are guidance-contract checks, not simulated access-control enforcement; no denied file was accessed. Subprocesses are timeout-bounded; absent Pester-v5 is explicitly skipped/not executed.
Ruby harness follow-up (
728658efd): The portable unittest command above was rerun with a 180s outer subprocess bound: exit 0, 17/17, no skips, including actual Pester9/9. Three new tests check selection/count guidance and execute the two published Bash blocks with synthetic runners: four executions, expected exits 0,0,23,23, preserved stdout summaries/stderr, and exactly one selected runner per invocation (30s bound each). GNU Bash5.3.15(2)-release on Git for Windows. Both unchanged upstream commands were separately reproduced returning Bash exit 0 despite runner exit 23. Eval-quality againstorigin/mainand whitespace check exit 0, existing warnings only. No eval YAML or framework/plugin routes changed in this follow-up. Ruby/Bundler are absent: synthetic exit propagation is not native RSpec/Rake/Minitest integration proof; no tools installed.Ruby/RSpec and Kotlin/Gradle/JVM tools are absent.
rustc --versionandcargo --versioneach exit 2,tool_not_installed; installed launchers are not a usable toolchain. Those three samples have structural evidence only. No follow-up tooling/global settings were installed or changed.All exit 0: plugin checks 23 skills/10 agents, 6 skills/1 agent, eval-quality with existing warnings only, whitespace clean. Both plugins/regression files were rerun after the reader correction; eval-quality/whitespace repeated after final Owner grading. Validator restore followed a proven missing-assets failure. Owning eval replaces deprecated
config:withdefaults:and adds capability/risk/journey tags; rubric weights, plugin metadata and external domains are unchanged.Native/reference contracts: Original xUnit fixtures ran 3/3,3/3,2/2 tests. Migrated goldens ran 3/3,2/2 on MSTest4.4.1, Test.Sdk18.9.0, SDK11.0.100-rc.1.26425.128, net8.0/VSTest. Inputs: xunit.v3 3.2.2, adapter 3.1.5, Test.Sdk 17.13.0. AwesomeAssertions 9.6.0, local/global/project imports and original assertions survive. Owner reflection requires one alice Owner on each of exactly three discoverable methods. Helper digests and complete staged input hashes are authenticated.
Correct states exit 0, with exact nonzero test counts/markers. Windows replay resolves specification
python3to installedpython. Final replay accepts all seven deterministic grader invocations against three reference states. Initial seven mutations reject lost/wrong/duplicate Owner, silent conflict rewriting and separate local/global/project FluentAssertions imports. Runner replay: 17 commands, package VSTest and MSTest.Sdk4.4.1 + UseVSTest=true accepted, SDK references run 3/2 tests, six native/bridged-MTP mutations and a boundary source edit rejected.Final assertion/inventory replay: 18 commands, arrow/block formatting accepted, four weakened-assertion/comment-decoy mutations rejected, both complete read-only goldens accepted, six add-file mutations rejected (Directory.Build.props, Directory.Build.targets and arbitrary JSON in each case). Inventory covers every file outside known
.eval/.gitmetadata andbin/obj, with exact supplied hashes, not extension filtering.Final Owner assertion guard:
verify_owner_assertions.pyexits 0 across nine native command logs. Original three assertion bodies are required alongside reflected names/Owner/discovery metadata. Golden and three individual block-format variants pass. Removing each assertion independently or emptying all three bodies produces four rejected mutations (exit1). The final seven-grader replay, production parser and full-PR eval-quality pass after the authenticated Owner digest update. Logs/scripts/reference fixtures are retained inmigration-review-verified,migration-review-final-graders,runner-review,final-eval-guards,owner-assertion-reviewand final quality logs.Production probe — executed, non-passing: Installed VallyCLI0.9.0, not claimed equivalent to documented CI0.14:
Lint/dry run exit 0 (one default-scoring warning). Actual run exits 1 after 8/8 executor records completed successfully, zero executor errors/timeouts: four stimuli × baseline/isolated × one run, claude-opus-4.6 executor/judge, normal 8 workers, declared 20m timeout, outer1320s bound. Runtime used
f5af583de; subsequent82da18db8/aa079b4c8grader strengthening was re-linted and deterministically replayed, not model-rerun.Authenticated command graders returned 9009 because Windows
python3resolves to an uninstalled Microsoft Store alias. That host prerequisite was classified before content changes. All three positive isolated trials invoked the target; advisory trial did not, but its rubric caught an unwanted MSTest recommendation. The judge also flagged a possible project-import omission, not certified by the blocked deterministic grader. Neither arm is green; no preference W/T/L/powered verdict is claimed. No plugin arm, second family, rerun or broad matrix. Owning suite: 16 preference stimuli + one dormant boundary; four-case probe is deliberately below power floor. Raw results/trajectories/metadata/run.log retained inproduction-review\results\2026-10-09T13-12-26-559Z, run ID8ef8bd30-2836-4785-a8cf-b482aa8770b7; Vally cleaned temporary model workspaces.Real selected collection: Selected three-test project references another project with an intentionally failing test.
dotnet test TestProject.csproj --logger "console;verbosity=normal"anddotnet test Selected.slnf --logger "console;verbosity=normal"both exit 0, run only the selected three tests, never the failing dependency. Published inventory agrees. Retainedverify_collection_graph.py/collection-graphcontain fixture/logs.Composed runner discovery/execution: Every command exits 0; discovery contains real
Invoice_is_discovered; direct/solution runs execute it.dotnet test Invoice.Tests.csproj --no-build;dotnet test Billing.sln --list-tests --no-build;dotnet test Billing.sln --no-builddotnet test Invoice.Tests.csproj --no-build;dotnet test Billing.sln --no-build -p:TestingPlatformCaptureOutput=false -- --list-tests;dotnet test Billing.sln --no-builddotnet test --project Invoice.Tests.csproj --no-build;dotnet test --solution Billing.slnx --list-tests --no-build;dotnet test --solution Billing.slnx --no-buildQualification: SDK10.0.401 also accepted positional native project syntax (
dotnet test Invoice.Tests.csproj --no-build). Guidance aligns documented named forms, not universal parser failure; the initial negative fixture expectation was corrected.Namespace/attribute probes: AwesomeAssertions9.6.0 with MSTest3.8.0/exact4.4.1.
dotnet run -v:qanddotnet run --project <retained-session-artifacts>\native-contracts\metadata\Metadata.csproj -v:qexit 0; invalid class Owner viadotnet build --no-restore -v:qfails CS0592.C++ native sample: Published blocks/minimal interface header compiled as C++20, MSVC19.51.36260.0, Catch2v3.8.1 commit
2b60af89e23d28eefc081bc930831ee9d45ea58b, after installed vcvars64.bat setup:All exit 0: 2 cases/9 sections/12 assertions, CTest1/1. NMake ignores parallel; incidental missing-vswhere warning did not prevent build/execution. Missing Catch2 was established before obtaining the pinned dependency. Fixture/log/XML evidence retained.
Original Pester enum failure,
.slnxmisselection and sibling leakage were reproduced before correction. These are local native/reference checks, not CI results or certified model-quality improvement. No broad Vally/model-family matrix was launched.Checklist
eng/known-domains.txtfor any new external domains referenced by skill content. (N/A: no new domains.)Evaluation changes only
For every eval-related change, including a new eval:
Three correctness scenarios and one uncued dormancy guard; existing already-MSTest no-op retained. Suite16 preference + one dormancy; four-case probe is not powered quality evidence.