You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Add .NET 11 Process API skill on a repository branch for evals - #1284
Promote @shaikhsharukh's process-api-net11 skill and eval contribution from #866 to a branch in dotnet/skills. All original commits, authorship, and merge history are preserved unchanged. The original PR remains open.
Why: The evaluation workflow policy disables secret-backed evals for fork PRs. A trusted repository branch permits those evaluations without changing the security rules.
Impact: Follow-up commits address Adam Sitnik's wording and platform comments, correct API type descriptions, document shell/redirection and inherited-handle restrictions, and guard the teardown setter on supported platforms, including Android. Eval graders accept valid capture and tuple forms, reject event calls rather than bare API names, and require actual line text. The latest commit, 7281bbee3, adds an inline ATIF golden response for each of the eight scenarios and a reproducible acceptance/mutation checker. References are judge evidence, not files supplied to the model's task environment.
Six distinct preference scenarios cover capture, teardown, tuple reading, streaming, independent launch, and handle inheritance. Two dormancy scenarios protect .NET 8 and .NET 10 routing. No workflow configuration or plugin metadata changes are included. Existing team ownership remains; #1290 proposes Adam's ownership separately.
python tests\dotnet11\process-api-net11\validate_references.py — passed: all eight ATIF golden responses satisfy their deterministic output graders; all eight realistic mutations fail their intended grader. The checker uses the production quality gate's ATIF and regex helpers.
python eng\eval-quality\check_eval_quality.py --base-ref origin/main — passed: "No errors." All eight target scenarios now contain golden references. The limited-power warning is acknowledged below.
git diff --check and git diff --cached --check — passed.
git merge-base --is-ancestor e1d59f9360eeaf41621f0d96c5804d04fff06d83 HEAD — passed, exit code 0; original-source history is preserved.
Session-only regression checks — passed all 70 output acceptance/mutation cases, including replay of three recorded false-failure answers.
dotnet run --project eng\skill-validator\src\SkillValidator.csproj -- check --plugin .\plugins\dotnet11 — previously could not start locally because the machine lacks the .NET 11 SDK required by global.json. No generated C# compile or runtime execution is claimed by the golden checker.
gh workflow run evaluation.yml --repo dotnet/skills --ref main -f pr_number=1284 -f head_sha=7281bbee346a66809c5d8884c1469fc40addbb52 -f matrix_profile=default — dispatched 37973791511 under the normal production profile. This final-head run is pending; it is not claimed to have passed.
The completed production run 37971249727 evaluated 64dd6348ab0cf99fa66ef1e003f7d699ec8d16bf, before golden references were added. Results are separate by executor, not pooled:
Executor
Wins / ties / losses
One-sided p-value
State
Dormancy
Claude Sonnet 5
6 / 0 / 0
0.015625
VALID_PASS
2 satisfied, 0 violated
GPT-5.6 Luna
3 / 3 / 0
0.125
VALID_NO_CHANGE
2 satisfied, 0 violated
Both results report zero unresolved comparison errors and underpowered: false. GPT's result does not establish a credible improvement. These results do not certify the final head. Earlier runs and their limits remain available in the PR's evaluation comments.
Checklist
I searched existing issues and pull requests to avoid duplicates.
I kept this pull request focused and avoided unrelated refactors.
I added or updated tests, evals, or documentation when changing skill or agent behavior.
I updated CODEOWNERS when adding or moving owned content.
I updated all marketplace manifests when plugin metadata changed.
I updated eng/known-domains.txt for any new external domains referenced by skill content.
Conditional items are addressed as follows: CODEOWNERS changes are not required here because the new skill and eval already fall under the existing dotnet11 reviewer team; adding Adam is scoped to #1290. Marketplace updates are not applicable because plugin metadata is unchanged. Domain updates are not applicable because no new external domain is referenced by skill content. Checked conditional items mean their applicability was reviewed, not that unnecessary files were changed.
Evaluation changes only
For every eval-related change, including a new eval:
Scenarios are necessary, distinct, and use natural prompts.
Graders cover the full result, accept the golden result, and reject a realistic mutation.
No-op, dormancy, and statistical power are covered where needed.
I ran the applicable production evaluation path and recorded the result above.
If this fixes a failed eval, I classified the failure before editing skill content.
If this broadly changes routing or behavior, I checked separate model-family evidence.
Scenario and grader evidence: Each prompt requests a C# program and project XML, and each golden response contains both. Deterministic checks cover the relevant target/API/output contracts; prompt graders cover configuration and behavioral correctness. All eight references pass the deterministic graders. One realistic mutation per scenario fails its intended grader, including missing stderr, disabled teardown, lost line text, null inheritance, and .NET 11 APIs on older targets. This is response validation, not compiled behavioral proof.
Restraint and power: A rewrite no-op case is not applicable because these tasks request printed examples, not edits to an existing project. Both older-framework dormancy contracts are retained and satisfied in the completed production run. Six independent preference scenarios meet the gate's eligibility floor; five discordant wins with no losses can reach p<=0.05. This portfolio has limited power when ties or losses occur, as the GPT result shows. No duplicate scenarios or repeated runs were added to manufacture significance, and high statistical power is not claimed.
Production and model-family evidence: The official production path completed with both executor families on 64dd6348a; its results and limitations are recorded above. The final-head rerun was dispatched with the normal profile and remains pending. These checked items record the work and available evidence, not a passing verdict for the final head.
Failure classification: Missing tags were a spec-structure defect. Bare API-name bans and the capture-redirection/tuple-style restrictions caused deterministic false failures. A judge's denial of a real .NET 11 API was stale knowledge, not a reason to replace the API. Those findings were checked before the focused fixes. Golden references now provide concrete expected responses; a pinned-SDK compile/behavior oracle remains a separate improvement.
Describe ProcessExitStatus and ProcessTextOutput as sealed classes and retain ProcessOutputLine as a readonly struct. Use property summaries instead of fictitious framework type declarations, and explain command-specific exit-code handling.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6
Require control-flow platform guard for KillOnParentExit
tests/dotnet11/process-api-net11/eval.yaml:57
This grader accepts both an unconditional KillOnParentExit = true and the current conditional RHS, even though neither control-flow-guards the platform-specific setter in a generic net11.0 sample. Require a supported-platform if guard that dominates the true assignment so a CA1416-producing answer cannot pass the deterministic contract.
Replace unverifiable skill activation rubric
tests/dotnet11/process-api-net11/eval.yaml:209
The prompt judge cannot observe whether a skill was loaded, and activation is already measured separately by expect_activation: false. Replace this skill-name rubric with an outcome the response can demonstrate.
This issue also appears on line 238 of the same file.
silent: true redirects standard input as well as output/error, and Unix uses /dev/null rather than Windows' NUL. Describe the platform null device so this cross-platform API guidance is accurate.
This issue also appears on line 76 of the same file.
Add ordered workflow and validation steps to the skill
The new skill has reference sections and examples but no ordered workflow or validation section. SkillProfiler therefore reports “No numbered workflow steps” (eng/skill-validator/src/Check/SkillProfiler.cs:263-264,417-418), and the repository quality bar requires skills to be actionable and verifiable (CONTRIBUTING.md:553-558). Add a short ordered decision/validation sequence covering API selection, required redirects/platform guards, build/run verification, and truthful reporting.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
process-api-net11
claude-sonnet-5
✅ Improved
n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded
🔴 0.60
—
Review overfit evidence.
process-api-net11
gpt-5.6-luna
➖ Improvement signal, unproven
n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 2 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Auto-teardown child processes on parent exit in .NET 11
Eligible
-100.0%
-100.0%
0/0/1
▼ Fire-and-forget process start in .NET 11
Eligible
-100.0%
-100.0%
0/0/1
= Non-activation: Running a process on .NET 8
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Non-activation: capture process output on .NET 10
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Auto-teardown child processes on parent exit in .NET 11: Response A provides a correct, working solution with code that compiles and runs, backed by research from actual Microsoft documentation. Response B's code references a non-existent API (Process.StartAndForget) that will fail at compile-time, making it non-functional despite h...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — process-api-net11 (claude-sonnet-5)
Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +80.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better
Repeated-run reliability (not used by the gate): 8 paired runs (7W/1T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Non-activation: Running a process on .NET 8
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Non-activation: Running a process on .NET 8: Both responses fully satisfy the requested output with effectively identical, valid minimal C# and .NET 8 project configuration. Their extra explanatory notes do not materially affect correctness.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37828027407 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b4906327ad00c5cc817c3d2e5b329e5310457ce7/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
👋 @AbhitejJohn — this PR has 7 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
The reason will be displayed to describe this comment to others. Learn more.
🔵 Needs a closer look
Platform guidance and several deterministic evaluation contracts remain incomplete.
0 open findings
Previously missed (7)
In code that hasn't changed since last review
Bind output checks to the iterated ProcessOutputLine
tests/dotnet11/process-api-net11/eval.yaml:287
These checks match any .ToString() and any .StandardError in the response, not members of the ProcessOutputLine yielded by ReadAllLinesAsync. For example, process.ToString() plus process.StandardError satisfies both while the loop never reads the line text or that line's source. Tie both checks to the await foreach element (or use a semantic/executable grader), and add this adversarial mutation.
Ban SafeProcessHandle.Start in the net8 guard
tests/dotnet11/process-api-net11/eval.yaml:484
The net8 guard still misses SafeProcessHandle.Start, although SKILL.md:21 explicitly makes that .NET 11 API a use case. An answer can include the required Process.Start and also recommend SafeProcessHandle.Start; all deterministic graders then pass despite containing an unavailable API. Add that invocation to the banned API alternation.
Ban SafeProcessHandle.Start in the net10 guard
tests/dotnet11/process-api-net11/eval.yaml:548
The net10 guard has the same gap for SafeProcessHandle.Start, which is part of this skill's routed .NET 11 surface. A response can satisfy the legacy capture checks and still append code using that unavailable API without failing any deterministic grader. Ban this call here as well.
The runtime-visible scope accepts every net11 target, but the .NET runtime marks Process.Run*, RunAndCaptureText*, and StartAndForget unsupported on iOS and tvOS. The body only mentions that restriction for SafeProcessHandle.Start, so routing can recommend unavailable APIs on those platforms. Add the mobile-platform boundary to the description and general usage guidance, or explicitly route such requests to an unsupported-platform answer.
Clarify StartDetached does not guarantee child survival
StartDetached does not unconditionally ensure that the child survives the parent: KillOnParentExit = true is valid alongside it and still terminates the child. State the detachment behavior without the absolute guarantee and call out the conflicting lifetime option.
This new skill provides reference material and examples but no ordered workflow or validation section. skill-validator will emit “No numbered workflow steps,” and agents are never told to inspect the TFM/platform or compile/probe generated code. Add a numbered workflow plus observable validation and truthful failure-reporting steps; CONTRIBUTING.md:553-579 requires skills to be verifiable and load-bearing API claims to be compiled or probed.
Remove unobservable skill-load target from the rubric
tests/dotnet11/process-api-net11/eval.yaml:489
expect_activation: false already provides the observable dormancy contract, while the answer judge cannot determine whether a skill was loaded. Naming the target in this rubric is therefore unobservable and overfit; remove it and retain only answer outcomes. The final production report also rates this eval's overfit as high.
This issue also appears on line 559 of the same file.
Regarding AI comments about iOS and tvOS like this:
The runtime-visible scope accepts every net11 target, but the .NET runtime marks Process.Run*, RunAndCaptureText*, and StartAndForget unsupported on iOS and tvOS.
Spawning processes is not supported on these platforms, using any APIs (old or new). This does not need to be repeated in the doc.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
process-api-net11
claude-sonnet-5
✅ Improved
n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded
🟡 0.42
—
Review overfit evidence.
process-api-net11
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 2 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Deadlock-free full-output read via instance ReadAllText in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
= Fire-and-forget process start in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
= Non-activation: Running a process on .NET 8
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Run and capture process output in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
= Stream tagged process output lines in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
= Whitelist inherited handles for child process in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Deadlock-free full-output read via instance ReadAllText in .NET 11: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — process-api-net11 (claude-sonnet-5)
Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better
Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Non-activation: Running a process on .NET 8
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Non-activation: Running a process on .NET 8: Both directly satisfy the requested code and project file. A is marginally stronger because its brief Windows applicability note is accurate, whereas B adds irrelevant and likely incorrect API-version commentary.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37980471884 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/95a0fd1a63ea93c1e270a47af1217d3629398e51/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Address Adam's review: keep general supported-platform scope without repeating the iOS and tvOS process-spawning limitation for an individual API. Preserve feature-specific platform guards.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6
Agreed. Removed the per-API iOS/tvOS note in 58ce0e2 and kept the general supported-platform wording. The feature-specific teardown guard is unchanged.
The reason will be displayed to describe this comment to others. Learn more.
🔵 Needs a closer look
Deterministic grader gaps and inaccurate lifecycle/concurrency guidance remain unresolved.
0 open findings
Previously missed (6)
In code that hasn't changed since last review
Require complete .csproj XML block for target framework
tests/dotnet11/process-api-net11/eval.yaml:55
This only proves that net11.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.
This issue also appears in the following locations of the same file:
line 131
line 194
line 275
line 348
line 419
Guard KillOnParentExit by supported platform
tests/dotnet11/process-api-net11/eval.yaml:128
An unconditional KillOnParentExit = true passes this deterministic contract even though the rubric and PR require the setter itself to be guarded for Windows, Linux, or Android. Add a deterministic guard check and a mutation that removes or bypasses the guard; otherwise the eval can accept the exact platform bug it claims to cover.
Bind output content to each ProcessOutputLine
tests/dotnet11/process-api-net11/eval.yaml:284
The unrestricted .ToString( alternative can be satisfied by any unrelated call, such as DateTime.Now.ToString(), while the loop never emits ProcessOutputLine text. Bind Content/ToString() to the await foreach variable so this grader actually proves that each line's text is consumed.
Require complete .csproj XML block for target framework
tests/dotnet11/process-api-net11/eval.yaml:478
This only proves that net8.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.
Require complete .csproj XML block for target framework
tests/dotnet11/process-api-net11/eval.yaml:545
This only proves that net10.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.
Clarify limited scope of concurrent process starts
This statement is inaccurate: the Windows implementation still uses a global process-start lock (a shared/read lock for explicit handle lists and a write lock otherwise). The supported benefit is concurrent starts with different explicit lists, not the absence of a global lock.
This issue also appears on line 131 of the same file.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
process-api-net11
claude-sonnet-5
✅ Improved
n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded
🔴 0.64
—
Review overfit evidence.
process-api-net11
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 2 dormancy excluded
🟡 0.40
Activation-only stop: plugin 1 failed run
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +55.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Warnings: Activation-only stop: plugin 1 failed run
Overfit: Moderate (score 0.40)
Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Auto-teardown child processes on parent exit in .NET 11
Eligible
+100.0%
+100.0%
1/0/0
= Fire-and-forget process start in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
= Non-activation: Running a process on .NET 8
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Whitelist inherited handles for child process in .NET 11
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Fire-and-forget process start in .NET 11: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — process-api-net11 (claude-sonnet-5)
Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better
Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Non-activation: Running a process on .NET 8
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Non-activation: Running a process on .NET 8: The requested code and project files are equivalently correct, but A's explanatory note correctly says the relevant default is UseShellExecute=false on modern .NET. B incorrectly claims the string overload launches via the shell, so A is slightly more accurate overall.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37999637321 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/58ce0e2dbf400c49cb951f5c0bf297709d884b35/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
✅ Approved by @adamsitnik. cc @dotnet/skills-merge-approvers — ready to merge.
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Promote @shaikhsharukh's
process-api-net11skill and eval contribution from #866 to a branch indotnet/skills. All original commits, authorship, and merge history are preserved unchanged. The original PR remains open.Why: The evaluation workflow policy disables secret-backed evals for fork PRs. A trusted repository branch permits those evaluations without changing the security rules.
Impact: Follow-up commits address Adam Sitnik's wording and platform comments, correct API type descriptions, document shell/redirection and inherited-handle restrictions, and guard the teardown setter on supported platforms, including Android. Eval graders accept valid capture and tuple forms, reject event calls rather than bare API names, and require actual line text. The latest commit,
7281bbee3, adds an inline ATIF golden response for each of the eight scenarios and a reproducible acceptance/mutation checker. References are judge evidence, not files supplied to the model's task environment.Six distinct preference scenarios cover capture, teardown, tuple reading, streaming, independent launch, and handle inheritance. Two dormancy scenarios protect .NET 8 and .NET 10 routing. No workflow configuration or plugin metadata changes are included. Existing team ownership remains; #1290 proposes Adam's ownership separately.
Related issue
Refs #649 and #866. Original contribution by @shaikhsharukh.
Validation
python tests\dotnet11\process-api-net11\validate_references.py— passed: all eight ATIF golden responses satisfy their deterministic output graders; all eight realistic mutations fail their intended grader. The checker uses the production quality gate's ATIF and regex helpers.python eng\eval-quality\check_eval_quality.py --base-ref origin/main— passed: "No errors." All eight target scenarios now contain golden references. The limited-power warning is acknowledged below.git diff --checkandgit diff --cached --check— passed.git merge-base --is-ancestor e1d59f9360eeaf41621f0d96c5804d04fff06d83 HEAD— passed, exit code 0; original-source history is preserved.dotnet run --project eng\skill-validator\src\SkillValidator.csproj -- check --plugin .\plugins\dotnet11— previously could not start locally because the machine lacks the .NET 11 SDK required byglobal.json. No generated C# compile or runtime execution is claimed by the golden checker.gh workflow run evaluation.yml --repo dotnet/skills --ref main -f pr_number=1284 -f head_sha=7281bbee346a66809c5d8884c1469fc40addbb52 -f matrix_profile=default— dispatched 37973791511 under the normal production profile. This final-head run is pending; it is not claimed to have passed.The completed production run 37971249727 evaluated
64dd6348ab0cf99fa66ef1e003f7d699ec8d16bf, before golden references were added. Results are separate by executor, not pooled:Both results report zero unresolved comparison errors and
underpowered: false. GPT's result does not establish a credible improvement. These results do not certify the final head. Earlier runs and their limits remain available in the PR's evaluation comments.Checklist
eng/known-domains.txtfor any new external domains referenced by skill content.Conditional items are addressed as follows: CODEOWNERS changes are not required here because the new skill and eval already fall under the existing dotnet11 reviewer team; adding Adam is scoped to #1290. Marketplace updates are not applicable because plugin metadata is unchanged. Domain updates are not applicable because no new external domain is referenced by skill content. Checked conditional items mean their applicability was reviewed, not that unnecessary files were changed.
Evaluation changes only
For every eval-related change, including a new eval:
Scenario and grader evidence: Each prompt requests a C# program and project XML, and each golden response contains both. Deterministic checks cover the relevant target/API/output contracts; prompt graders cover configuration and behavioral correctness. All eight references pass the deterministic graders. One realistic mutation per scenario fails its intended grader, including missing stderr, disabled teardown, lost line text, null inheritance, and .NET 11 APIs on older targets. This is response validation, not compiled behavioral proof.
Restraint and power: A rewrite no-op case is not applicable because these tasks request printed examples, not edits to an existing project. Both older-framework dormancy contracts are retained and satisfied in the completed production run. Six independent preference scenarios meet the gate's eligibility floor; five discordant wins with no losses can reach p<=0.05. This portfolio has limited power when ties or losses occur, as the GPT result shows. No duplicate scenarios or repeated runs were added to manufacture significance, and high statistical power is not claimed.
Production and model-family evidence: The official production path completed with both executor families on
64dd6348a; its results and limitations are recorded above. The final-head rerun was dispatched with the normal profile and remains pending. These checked items record the work and available evidence, not a passing verdict for the final head.Failure classification: Missing tags were a spec-structure defect. Bare API-name bans and the capture-redirection/tuple-style restrictions caused deterministic false failures. A judge's denial of a real .NET 11 API was stale knowledge, not a reason to replace the API. Those findings were checked before the focused fixes. Golden references now provide concrete expected responses; a pinned-SDK compile/behavior oracle remains a separate improvement.