Skip to content

Add .NET 11 Process API skill on a repository branch for evals - #1284

Open
AbhitejJohn wants to merge 23 commits into
mainfrom
abhitejjohn-skill-eval-follow-up
Open

AbhitejJohn wants to merge 23 commits into
mainfrom
abhitejjohn-skill-eval-follow-up

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Promote @shaikhsharukh's process-api-net11 skill and eval contribution from #866 to a branch in dotnet/skills. All original commits, authorship, and merge history are preserved unchanged. The original PR remains open.

Why: The evaluation workflow policy disables secret-backed evals for fork PRs. A trusted repository branch permits those evaluations without changing the security rules.

Impact: Follow-up commits address Adam Sitnik's wording and platform comments, correct API type descriptions, document shell/redirection and inherited-handle restrictions, and guard the teardown setter on supported platforms, including Android. Eval graders accept valid capture and tuple forms, reject event calls rather than bare API names, and require actual line text. The latest commit, 7281bbee3, adds an inline ATIF golden response for each of the eight scenarios and a reproducible acceptance/mutation checker. References are judge evidence, not files supplied to the model's task environment.

Six distinct preference scenarios cover capture, teardown, tuple reading, streaming, independent launch, and handle inheritance. Two dormancy scenarios protect .NET 8 and .NET 10 routing. No workflow configuration or plugin metadata changes are included. Existing team ownership remains; #1290 proposes Adam's ownership separately.

Related issue

Refs #649 and #866. Original contribution by @shaikhsharukh.

Validation

  • python tests\dotnet11\process-api-net11\validate_references.py — passed: all eight ATIF golden responses satisfy their deterministic output graders; all eight realistic mutations fail their intended grader. The checker uses the production quality gate's ATIF and regex helpers.
  • python eng\eval-quality\check_eval_quality.py --base-ref origin/main — passed: "No errors." All eight target scenarios now contain golden references. The limited-power warning is acknowledged below.
  • git diff --check and git diff --cached --check — passed.
  • git merge-base --is-ancestor e1d59f9360eeaf41621f0d96c5804d04fff06d83 HEAD — passed, exit code 0; original-source history is preserved.
  • Session-only regression checks — passed all 70 output acceptance/mutation cases, including replay of three recorded false-failure answers.
  • dotnet run --project eng\skill-validator\src\SkillValidator.csproj -- check --plugin .\plugins\dotnet11 — previously could not start locally because the machine lacks the .NET 11 SDK required by global.json. No generated C# compile or runtime execution is claimed by the golden checker.
  • gh workflow run evaluation.yml --repo dotnet/skills --ref main -f pr_number=1284 -f head_sha=7281bbee346a66809c5d8884c1469fc40addbb52 -f matrix_profile=default — dispatched 37973791511 under the normal production profile. This final-head run is pending; it is not claimed to have passed.

The completed production run 37971249727 evaluated 64dd6348ab0cf99fa66ef1e003f7d699ec8d16bf, before golden references were added. Results are separate by executor, not pooled:

Executor Wins / ties / losses One-sided p-value State Dormancy
Claude Sonnet 5 6 / 0 / 0 0.015625 VALID_PASS 2 satisfied, 0 violated
GPT-5.6 Luna 3 / 3 / 0 0.125 VALID_NO_CHANGE 2 satisfied, 0 violated

Both results report zero unresolved comparison errors and underpowered: false. GPT's result does not establish a credible improvement. These results do not certify the final head. Earlier runs and their limits remain available in the PR's evaluation comments.

Checklist

  • I searched existing issues and pull requests to avoid duplicates.
  • I kept this pull request focused and avoided unrelated refactors.
  • I added or updated tests, evals, or documentation when changing skill or agent behavior.
  • I updated CODEOWNERS when adding or moving owned content.
  • I updated all marketplace manifests when plugin metadata changed.
  • I updated eng/known-domains.txt for any new external domains referenced by skill content.

Conditional items are addressed as follows: CODEOWNERS changes are not required here because the new skill and eval already fall under the existing dotnet11 reviewer team; adding Adam is scoped to #1290. Marketplace updates are not applicable because plugin metadata is unchanged. Domain updates are not applicable because no new external domain is referenced by skill content. Checked conditional items mean their applicability was reviewed, not that unnecessary files were changed.

Evaluation changes only

For every eval-related change, including a new eval:

  • Scenarios are necessary, distinct, and use natural prompts.
  • Graders cover the full result, accept the golden result, and reject a realistic mutation.
  • No-op, dormancy, and statistical power are covered where needed.
  • I ran the applicable production evaluation path and recorded the result above.
  • If this fixes a failed eval, I classified the failure before editing skill content.
  • If this broadly changes routing or behavior, I checked separate model-family evidence.

Scenario and grader evidence: Each prompt requests a C# program and project XML, and each golden response contains both. Deterministic checks cover the relevant target/API/output contracts; prompt graders cover configuration and behavioral correctness. All eight references pass the deterministic graders. One realistic mutation per scenario fails its intended grader, including missing stderr, disabled teardown, lost line text, null inheritance, and .NET 11 APIs on older targets. This is response validation, not compiled behavioral proof.

Restraint and power: A rewrite no-op case is not applicable because these tasks request printed examples, not edits to an existing project. Both older-framework dormancy contracts are retained and satisfied in the completed production run. Six independent preference scenarios meet the gate's eligibility floor; five discordant wins with no losses can reach p<=0.05. This portfolio has limited power when ties or losses occur, as the GPT result shows. No duplicate scenarios or repeated runs were added to manufacture significance, and high statistical power is not claimed.

Production and model-family evidence: The official production path completed with both executor families on 64dd6348a; its results and limitations are recorded above. The final-head rerun was dispatched with the normal profile and remains pending. These checked items record the work and available evidence, not a passing verdict for the final head.

Failure classification: Missing tags were a spec-structure defect. Bare API-name bans and the capture-redirection/tuple-style restrictions caused deterministic false failures. A judge's denial of a real .NET 11 API was stale knowledge, not a reason to replace the API. Those findings were checked before the focused fixes. Golden references now provide concrete expected responses; a pinned-SDK compile/behavior oracle remains a separate improvement.

shaikhsharukh and others added 17 commits July 7, 2026 20:13
…te struct definitions, API notes, examples, and migrate eval to Vally schema
Address Adam Sitnik's latest wording and supported-platform feedback. Add required eval result-slice tags and accept platform guards in the child-process grader.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6
@AbhitejJohn
AbhitejJohn requested a review from a team as a code owner October 8, 2026 18:49
Copilot AI balanced review requested due to automatic review settings October 8, 2026 18:49
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
❌ dotnet11 process-api-net11 0/4 0%
✅ dotnet11 system-text-json-net11 15/16 93.8%
Uncovered: dotnet11/process-api-net11
  • [CodePattern] [Err] (line 191)
  • [CodePattern] [Out] (line 191)
  • [CodePattern] ValueTuple (line 91)
  • [CodePattern] CancellationToken (line 57)
Uncovered: dotnet11/system-text-json-net11
  • [CodePattern] [guid] (line 48)

Describe ProcessExitStatus and ProcessTextOutput as sealed classes and retain ProcessOutputLine as a readonly struct. Use property summaries instead of fictitious framework type declarations, and explain command-specific exit-code handling.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The skill contains stale API definitions and platform/concurrency guidance, while several deterministic graders accept incorrect answers.

8 open findings
What changed in this PR

Promotes the .NET 11 Process API skill to a trusted repository branch so maintainers can run secret-backed evaluations.

Changes:

  • Adds guidance and examples for new Process APIs.
  • Adds six preference and two framework-dormancy eval cases.
  • Registers the skill in the .NET 11 plugin README.
File Description
plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md Adds Process API guidance and examples.
tests/​dotnet11/​process-api-net11/​eval.yaml Adds capability and routing evaluations.
plugins/​dotnet11/​README.md Lists the new skill.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread plugins/dotnet11/skills/process-api-net11/SKILL.md Outdated
Comment thread plugins/dotnet11/skills/process-api-net11/SKILL.md Outdated
Comment thread plugins/dotnet11/skills/process-api-net11/SKILL.md Outdated
Comment thread tests/dotnet11/process-api-net11/eval.yaml
Comment thread tests/dotnet11/process-api-net11/eval.yaml Outdated
Comment thread tests/dotnet11/process-api-net11/eval.yaml Outdated
Comment thread tests/dotnet11/process-api-net11/eval.yaml Outdated
Comment thread plugins/dotnet11/skills/process-api-net11/SKILL.md
Copilot AI balanced review requested due to automatic review settings October 8, 2026 18:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Platform safety, API documentation, skill structure, and deterministic eval coverage issues remain unresolved.

7 open findings
1 resolved since last review
Previously missed (4)

In code that hasn't changed since last review

Medium severity Require control-flow platform guard for KillOnParentExit

tests/​dotnet11/​process-api-net11/​eval.yaml:57

This grader accepts both an unconditional KillOnParentExit = true and the current conditional RHS, even though neither control-flow-guards the platform-specific setter in a generic net11.0 sample. Require a supported-platform if guard that dominates the true assignment so a CA1416-producing answer cannot pass the deterministic contract.

Medium severity Replace unverifiable skill activation rubric

tests/​dotnet11/​process-api-net11/​eval.yaml:209

The prompt judge cannot observe whether a skill was loaded, and activation is already measured separately by expect_activation: false. Replace this skill-name rubric with an outcome the response can demonstrate.

This issue also appears on line 238 of the same file.

Low severity Clarify silent mode's platform-specific null device behavior

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:54

silent: true redirects standard input as well as output/error, and Unix uses /dev/null rather than Windows' NUL. Describe the platform null device so this cross-platform API guidance is accurate.

This issue also appears on line 76 of the same file.

Low severity Add ordered workflow and validation steps to the skill

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:127

The new skill has reference sections and examples but no ordered workflow or validation section. SkillProfiler therefore reports “No numbered workflow steps” (eng/skill-validator/src/Check/SkillProfiler.cs:263-264,417-418), and the repository quality bar requires skills to be actionable and verifiable (CONTRIBUTING.md:553-558). Add a short ordered decision/validation sequence covering API selection, required redirects/platform guards, build/run verification, and truthful reporting.

🧠 Review effort: Balanced

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 1 improved, ➖ 1 result without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit b4906327ad00c5cc817c3d2e5b329e5310457ce7; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
process-api-net11 claude-sonnet-5 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded 🔴 0.60 — Review overfit evidence.
process-api-net11 gpt-5.6-luna ➖ Improvement signal, unproven n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 2 dormancy excluded 🟡 0.31 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, unproven — process-api-net11 (gpt-5.6-luna)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 2 dormancy excluded

Warnings: Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Auto-teardown child processes on parent exit in .NET 11 Eligible -100.0% -100.0% 0/0/1
▼ Fire-and-forget process start in .NET 11 Eligible -100.0% -100.0% 0/0/1
= Non-activation: Running a process on .NET 8 Excluded (activation contract) +0.0% +0.0% 0/1/0
= Non-activation: capture process output on .NET 10 Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Auto-teardown child processes on parent exit in .NET 11: Response A provides a correct, working solution with code that compiles and runs, backed by research from actual Microsoft documentation. Response B's code references a non-existent API (Process.StartAndForget) that will fail at compile-time, making it non-functional despite h...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — process-api-net11 (claude-sonnet-5)

Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +80.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded

Overfit: High (score 0.60)

Repeated-run reliability (not used by the gate): 8 paired runs (7W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Non-activation: Running a process on .NET 8 Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Non-activation: Running a process on .NET 8: Both responses fully satisfy the requested output with effectively identical, valid minimal C# and .NET 8 project configuration. Their extra explanatory notes do not materially affect correctness.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37828027407 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b4906327ad00c5cc817c3d2e5b329e5310457ce7/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 8, 2026
@github-actions github-actions Bot added the waiting-on-author PR state label label Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has 7 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Apply Adam's concise shell-default wording. Restrict SafeProcessHandle.Start guidance to supported platforms, excluding iOS and tvOS.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6
Copilot AI balanced review requested due to automatic review settings October 9, 2026 19:17

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Platform guidance and several deterministic evaluation contracts remain incomplete.

0 open findings

Previously missed (7)

In code that hasn't changed since last review

Medium severity Bind output checks to the iterated ProcessOutputLine

tests/​dotnet11/​process-api-net11/​eval.yaml:287

These checks match any .ToString() and any .StandardError in the response, not members of the ProcessOutputLine yielded by ReadAllLinesAsync. For example, process.ToString() plus process.StandardError satisfies both while the loop never reads the line text or that line's source. Tie both checks to the await foreach element (or use a semantic/executable grader), and add this adversarial mutation.

Medium severity Ban SafeProcessHandle.Start in the net8 guard

tests/​dotnet11/​process-api-net11/​eval.yaml:484

The net8 guard still misses SafeProcessHandle.Start, although SKILL.md:21 explicitly makes that .NET 11 API a use case. An answer can include the required Process.Start and also recommend SafeProcessHandle.Start; all deterministic graders then pass despite containing an unavailable API. Add that invocation to the banned API alternation.

Medium severity Ban SafeProcessHandle.Start in the net10 guard

tests/​dotnet11/​process-api-net11/​eval.yaml:548

The net10 guard has the same gap for SafeProcessHandle.Start, which is part of this skill's routed .NET 11 surface. A response can satisfy the legacy capture checks and still append code using that unavailable API without failing any deterministic grader. Ban this call here as well.

Low severity Document mobile platform limits for Process APIs

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:8

The runtime-visible scope accepts every net11 target, but the .NET runtime marks Process.Run*, RunAndCaptureText*, and StartAndForget unsupported on iOS and tvOS. The body only mentions that restriction for SafeProcessHandle.Start, so routing can recommend unavailable APIs on those platforms. Add the mobile-platform boundary to the description and general usage guidance, or explicitly route such requests to an unsupported-platform answer.

Low severity Clarify StartDetached does not guarantee child survival

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:131

StartDetached does not unconditionally ensure that the child survives the parent: KillOnParentExit = true is valid alongside it and still terminates the child. State the detachment behavior without the absolute guarantee and call out the conflicting lifetime option.

Low severity Add workflow and validation steps to the skill

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:140

This new skill provides reference material and examples but no ordered workflow or validation section. skill-validator will emit “No numbered workflow steps,” and agents are never told to inspect the TFM/platform or compile/probe generated code. Add a numbered workflow plus observable validation and truthful failure-reporting steps; CONTRIBUTING.md:553-579 requires skills to be verifiable and load-bearing API claims to be compiled or probed.

Low severity Remove unobservable skill-load target from the rubric

tests/​dotnet11/​process-api-net11/​eval.yaml:489

expect_activation: false already provides the observable dormancy contract, while the answer judge cannot determine whether a skill was loaded. Naming the target in this rubric is therefore unobservable and overfit; remove it and retain only answer outcomes. The final production report also rates this eval's overfit as high.

This issue also appears on line 559 of the same file.

🧠 Review effort: Balanced

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 9, 2026
@adamsitnik

Copy link
Copy Markdown
Member

Regarding AI comments about iOS and tvOS like this:

The runtime-visible scope accepts every net11 target, but the .NET runtime marks Process.Run*, RunAndCaptureText*, and StartAndForget unsupported on iOS and tvOS.

Spawning processes is not supported on these platforms, using any APIs (old or new). This does not need to be repeated in the doc.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 1 improved, ➖ 1 result without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 95a0fd1a63ea93c1e270a47af1217d3629398e51; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
process-api-net11 claude-sonnet-5 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded 🟡 0.42 — Review overfit evidence.
process-api-net11 gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 2 dormancy excluded 🟡 0.38 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, tie-limited — process-api-net11 (gpt-5.6-luna)

Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 2 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Deadlock-free full-output read via instance ReadAllText in .NET 11 Eligible +0.0% +0.0% 0/1/0
= Fire-and-forget process start in .NET 11 Eligible +0.0% +0.0% 0/1/0
= Non-activation: Running a process on .NET 8 Excluded (activation contract) +0.0% +0.0% 0/1/0
= Run and capture process output in .NET 11 Eligible +0.0% +0.0% 0/1/0
= Stream tagged process output lines in .NET 11 Eligible +0.0% +0.0% 0/1/0
= Whitelist inherited handles for child process in .NET 11 Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Deadlock-free full-output read via instance ReadAllText in .NET 11: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — process-api-net11 (claude-sonnet-5)

Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Non-activation: Running a process on .NET 8 Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Non-activation: Running a process on .NET 8: Both directly satisfy the requested code and project file. A is marginally stronger because its brief Windows applicability note is accurate, whereas B adds irrelevant and likely incorrect API-version commentary.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37980471884 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/95a0fd1a63ea93c1e270a47af1217d3629398e51/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 9, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 9, 2026
Address Adam's review: keep general supported-platform scope without repeating the iOS and tvOS process-spawning limitation for an individual API. Preserve feature-specific platform guards.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b7932b85-307e-41df-bd20-3ed764830ed6
Copilot AI balanced review requested due to automatic review settings October 9, 2026 21:42
@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

This does not need to be repeated in the doc.

Agreed. Removed the per-API iOS/tvOS note in 58ce0e2 and kept the general supported-platform wording. The feature-specific teardown guard is unchanged.

(Copilot, commenting on Abhitej's behalf.)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Deterministic grader gaps and inaccurate lifecycle/concurrency guidance remain unresolved.

0 open findings

Previously missed (6)

In code that hasn't changed since last review

Medium severity Require complete .csproj XML block for target framework

tests/​dotnet11/​process-api-net11/​eval.yaml:55

This only proves that net11.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.

This issue also appears in the following locations of the same file:

  • line 131
  • line 194
  • line 275
  • line 348
  • line 419
Medium severity Guard KillOnParentExit by supported platform

tests/​dotnet11/​process-api-net11/​eval.yaml:128

An unconditional KillOnParentExit = true passes this deterministic contract even though the rubric and PR require the setter itself to be guarded for Windows, Linux, or Android. Add a deterministic guard check and a mutation that removes or bypasses the guard; otherwise the eval can accept the exact platform bug it claims to cover.

Medium severity Bind output content to each ProcessOutputLine

tests/​dotnet11/​process-api-net11/​eval.yaml:284

The unrestricted .ToString( alternative can be satisfied by any unrelated call, such as DateTime.Now.ToString(), while the loop never emits ProcessOutputLine text. Bind Content/ToString() to the await foreach variable so this grader actually proves that each line's text is consumed.

Medium severity Require complete .csproj XML block for target framework

tests/​dotnet11/​process-api-net11/​eval.yaml:478

This only proves that net8.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.

Medium severity Require complete .csproj XML block for target framework

tests/​dotnet11/​process-api-net11/​eval.yaml:545

This only proves that net10.0 appears somewhere in the response, so an answer can omit the explicitly requested .csproj and still pass every deterministic grader. Match a complete project XML block containing the target framework instead.

Low severity Clarify limited scope of concurrent process starts

plugins/​dotnet11/​skills/​process-api-net11/​SKILL.md:121

This statement is inaccurate: the Windows implementation still uses a global process-start lock (a shared/read lock for explicit handle lists and a write lock otherwise). The supported benefit is concurrent starts with different explicit lists, not the absence of a global lock.

This issue also appears on line 131 of the same file.

🧠 Review effort: Balanced

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-review PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 9, 2026
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 1 improved, ➖ 1 result without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 58ce0e2dbf400c49cb951f5c0bf297709d884b35; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
process-api-net11 claude-sonnet-5 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded 🔴 0.64 — Review overfit evidence.
process-api-net11 gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 2 dormancy excluded 🟡 0.40 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, tie-limited — process-api-net11 (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +55.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 2 dormancy excluded

Warnings: Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.40)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Auto-teardown child processes on parent exit in .NET 11 Eligible +100.0% +100.0% 1/0/0
= Fire-and-forget process start in .NET 11 Eligible +0.0% +0.0% 0/1/0
= Non-activation: Running a process on .NET 8 Excluded (activation contract) +0.0% +0.0% 0/1/0
= Whitelist inherited handles for child process in .NET 11 Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Fire-and-forget process start in .NET 11: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — process-api-net11 (claude-sonnet-5)

Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded

Overfit: High (score 0.64)

Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Non-activation: Running a process on .NET 8 Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Non-activation: Running a process on .NET 8: The requested code and project files are equivalently correct, but A's explanatory note correctly says the relevant default is UseShellExecute=false on modern .NET. B incorrectly claims the string overload launches via the shell, so A is slightly more accurate overall.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1284 in dotnet/skills, download eval artifacts with gh run download 37999637321 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/58ce0e2dbf400c49cb951f5c0bf297709d884b35/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 9, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 9, 2026
@github-actions github-actions Bot added ready-to-merge PR state label and removed waiting-on-review PR state label labels Oct 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

✅ Approved by @adamsitnik. cc @dotnet/skills-merge-approvers — ready to merge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-to-merge PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants