Skip to content

Fix the GPU CI tiers and split run_tests.sh into tests/lib - #181

Open
limou102 wants to merge 4 commits into
mainfrom
fix/ci-runner-env-and-run-tests-refactor
Open

limou102 wants to merge 4 commits into
mainfrom
fix/ci-runner-env-and-run-tests-refactor

Conversation

@limou102

@limou102 limou102 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Why the engine / e2e tiers kept failing

  1. Runners on crsuse2-slog-003/004 cannot start a job. Those login hosts now cap each user at 256 MB pressure / 512 MB hard. Two idle Runner.Listeners already sit near 290 MB, so a fresh Runner.Worker is throttled past the listener's 30 s handshake, gets killed (exit 137), and GitHub reports "The self-hosted runner lost communication with the server" 10 minutes later. Every GPU job that landed there failed this way. Not fixable in this repo: those four runners stay in the crusoe pool, and jobs landing on them will keep failing this way until the cluster admins restore the hosts.
  2. Stale scheduler env on the 005 runners. The listeners had been up since 2026-09-21 and still carried the old SPUR_CONTROLLER_ADDR (three IPs). After the controller moved, every srun/sbatch failed with The request does not have valid authentication credentials ... refusing to re-forward an already-forwarded request, which the dispatcher retried as a transient error until its five submissions were spent.
  3. The disagg hold re-walked refused QoS rungs. amd-primus-cicd-qos and amd-primus-qos answer QOSGrpNodeLimit in every recent run, but each re-picked pair started the ladder from rung 0 again, spending two of the five submissions per pair.
  4. A stale engine test. tests/engine/sglang/test_wait_for_decode_args.py imported _wait_for_decode_until_stop (renamed to _run_startup_barrier_until_stop in de69b2b) and built EngineDeath(exit_status=...), so the file failed to collect.

Changes

  • .github/scripts/refresh_spur_env.sh + a step in engine, e2e-mixed and e2e-disag: export the host's current SPUR_* into $GITHUB_ENV, so both the tests and the reclaim step talk to the live controller.
  • The account/QoS rung is shared by every submission in a run; a scheduler auth refusal now stops the tier immediately and names the cause.
  • tests/run_tests.sh is split into tests/lib/{env,slurm,disagg,tiers}.sh plus container_pytest.sh (the in-container pytest that used to be inlined as quoted strings). CLI and env knobs are unchanged, and log lines are unchanged except the shortened "submitted to SLURM" banner and the unwritable-TMPDIR hint; comments are trimmed to at most two lines each. A dispatched tier now copies the remote leg's failure lines into the local summary instead of printing "no per-test detail was captured".
  • Fixed the stale sglang engine test.

Testing

  • tests/unit/e2e_harness mocked-scheduler tests (test_slurm_retry_limit, test_site_profiles, test_arch_overlay): all 69 pass on the refactored script.
  • Real cluster, CI env: tests/run_tests.sh engine dispatched, built both images, and passed (vllm all files; sglang after the test fix). On this PR's CI, every GPU job that ran on a 005 runner passed: engine, e2e-mixed (sglang), e2e-disag (vllm). The ones on 004 died with "runner lost communication".
  • Real cluster, manual run (no CI variables): e2e vllm mixed and engine were both accepted by SLURM (worker logs in node-local /tmp, output streamed back) and then cancelled; the INT/TERM trap scancelled the job each time.

lxgsbqylbk and others added 2 commits October 8, 2026 03:36
The barrier helper was renamed to _run_startup_barrier_until_stop and now wraps
both PD checks, and EngineDeath derives exit_status from observed/returncode.
The test still imported the old name, so the whole file failed to collect and
the engine tier went red on every run that reached a GPU node.

Signed-off-by: LI MOU <lxglbk@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Runner listeners outlive scheduler config changes: the 005 runners still held
the old SPUR_CONTROLLER_ADDR, so every srun/sbatch was refused with an auth
error that the dispatcher retried as transient. Each GPU job now re-reads the
host's SPUR_* settings (.github/scripts/refresh_spur_env.sh), and an auth
refusal stops the tier at once with the cause named.

The account/QoS rung is now shared by every submission in a run. Re-walking
the always-full top rungs for each re-picked disagg pair spent two of the five
submissions per pair, so one slow burst hold was enough to give up.

run_tests.sh keeps its CLI and messages; the logic moves to tests/lib (env,
slurm, disagg, tiers, and an in-container pytest helper), and a dispatched
tier now carries the remote leg's failure lines into the local summary.

Signed-off-by: LI MOU <lxglbk@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Copilot AI balanced review requested due to automatic review settings October 8, 2026 03:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The scheduler refresh script can expose host-provided credential values in GitHub Actions logs.

1 open finding
What changed in this PR

Improves GPU CI reliability and modularizes the test harness.

Changes:

  • Refreshes scheduler configuration and improves SLURM retry handling.
  • Splits run_tests.sh into focused library scripts.
  • Repairs stale SGLang tests and adds scheduler regressions.
File Description
.github/​workflows/​ci.yml Refreshes scheduler variables before GPU jobs.
.github/​scripts/​refresh_spur_env.sh Loads current host scheduler configuration.
tests/​run_tests.sh Becomes the modular harness entry point.
tests/​lib/​env.sh Provides environment and scratch setup.
tests/​lib/​slurm.sh Implements scheduling and retry behavior.
tests/​lib/​disagg.sh Implements disaggregated test orchestration.
tests/​lib/​tiers.sh Defines unit, engine, and mixed tiers.
tests/​lib/​container_pytest.sh Runs pytest inside containers.
tests/​engine/​sglang/​test_wait_for_decode_args.py Updates tests for renamed startup APIs.
tests/​unit/​e2e_harness/​test_slurm_retry_limit.py Adds scheduler retry regressions.
tests/​unit/​e2e_harness/​test_site_profiles.py Scans modular harness readers.
tests/​unit/​e2e_harness/​test_arch_overlay.py Validates image references across libraries.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

for f in /etc/environment /etc/profile.d/spur.sh; do
[ -r "$f" ] || continue
sed -nE 's/^[[:space:]]*(export[[:space:]]+)?(SPUR_[A-Z_]+)=["'\'']?([^"'\'']*)["'\'']?[[:space:]]*$/\2=\3/p' "$f"
done | awk -F= '!seen[$1]++' | tee -a "${GITHUB_ENV:?not running under GitHub Actions}"
lxgsbqylbk and others added 2 commits October 8, 2026 08:32
Signed-off-by: LI MOU <lxglbk@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants