Skip to content

[AMD][MI35X] Bump Qwen3.5 MXFP4 MI355X SGLang AgentX to v0.5.18 and retune all-reduce, prefill, and CUDA graph - #2737

Open
yichiche wants to merge 4 commits into
mainfrom
amd/qwen3.5-fp4-mi355x-sglang-agentic-v0.5.18
Open

[AMD][MI35X] Bump Qwen3.5 MXFP4 MI355X SGLang AgentX to v0.5.18 and retune all-reduce, prefill, and CUDA graph#2737
yichiche wants to merge 4 commits into
mainfrom
amd/qwen3.5-fp4-mi355x-sglang-agentic-v0.5.18

Conversation

@yichiche

@yichiche yichiche commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

Motivation

qwen3.5-fp4-mi355x-sglang-agentic-mtp is still on lmsysorg/sglang-rocm:v0.5.17-rocm720-mi35x-20260818, and three of its launch knobs are out of step with the sibling recipes on the same cluster.

First, all-reduce. The script correctly omits --enable-aiter-allreduce-fusion (removed in #2562 for TP2/EP2 EAGLE rank consistency) but never sets ROCM_QUICK_REDUCE_QUANTIZATION, so multi-GPU collectives run unquantized custom all-reduce. The published SGLang cookbook recipe for MXFP4 on MI355X calls for INT8-quantized quick all-reduce.

Second, the prefill budget is double the B200 sibling: --max-prefill-tokens 32768 / --chunked-prefill-size 32768 here versus 16384 / 16384 in benchmarks/single_node/agentic/qwen3.5_fp4_b200_sglang_mtp.sh.

Third, the decode CUDA graph is undersized for the batch the scheduler actually builds. --max-running-requests is 2*CONC, but the graph was only captured to min(CONC, 64), so every decode batch above CONC fell onto the eager path.

Modifications

Bump image on qwen3.5-fp4-mi355x-sglang-agentic-mtp to lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260825. The model, runner, and search space are untouched.

In benchmarks/single_node/agentic/qwen3.5_fp4_mi355x_sglang_mtp.sh, add export ROCM_QUICK_REDUCE_QUANTIZATION=INT8 alongside the existing aiter exports. The two all-reduce paths are mutually exclusive in SGLang: the AITER fused AR+RMSNorm path is gated on --enable-aiter-allreduce-fusion, and only with it off do collectives fall back to custom all-reduce, where the quick-reduce regime is read. Since this arm already omits the fusion flag, the env takes effect with no other launch change.

Lower --max-prefill-tokens and --chunked-prefill-size from 32768 to 16384, matching the B200 sibling.

Capture the decode CUDA graph to min(2*CONC, 128) instead of min(CONC, 64), so it covers --max-running-requests. The expression and the 128 cap are taken verbatim from the sibling MI355X AgentX recipe benchmarks/single_node/agentic/dsv4_fp4_mi355x_sglang_mtp.sh, which already runs this idiom on cluster:mi355x-amds.

Append the corresponding perf-changelog.yaml trigger.

Accuracy Tests

No accuracy-affecting logic changes in this repo. INT8 quick all-reduce quantizes the collective payload, so it is not bit-identical to the unquantized path; the AgentX eval rows continue to run real target-model verification and will show whether that matters. The prefill and CUDA-graph knobs change scheduling and kernel launch shape, not what is computed.

Benchmarking

Repo validation was run locally:

  • python3 -m pytest utils/matrix_logic/ -q → 232 passed.
  • bash -n benchmarks/single_node/agentic/qwen3.5_fp4_mi355x_sglang_mtp.sh → clean.
  • python3 utils/matrix_logic/generate_sweep_configs.py full-sweep --config-files configs/amd-master.yaml --model-prefix qwen3.5 --precision fp4 --scenario-type agentic-coding → 16 configs, all on v0.5.18-rocm720-mi35x-20260825.
  • Hand-evaluated the new CUDA-graph expression across the config's conc-list (TP2 1,4,8,12,16,20; TP4 1,4,8,12,16,20,24,28,32,40): it yields 2*CONC at every point, so the 128 cap never binds in the current sweep.

End-to-end MI355X AgentX numbers will come from the sweep triggered on this PR (full-sweep-fail-fast). Three knobs plus an image move together here, so the resulting numbers are not attributable to any single one of them; if the arm regresses, the all-reduce regime is the first knob to isolate.

Conflicts

#2693 touches the same script, the same config block, and also appends to perf-changelog.yaml, so the changelog append will collide and whichever lands second needs a rebase. The changes themselves are complementary.

Two notes for whoever reviews alongside #2693. That PR raises the TP4 concurrency ceiling to 64 via its HiCache arms — at CONC=64 the new min(2*CONC, 128) lands exactly on the 128 cap, which is the intended behaviour but worth knowing. And #2693's stated motivation is MI355X/B200 comparability: the B200 sibling still captures to min(CONC, 64), so this PR re-introduces a difference on that knob. It is harness tuning rather than a deployment-defining server arg, but it is a real difference and should not be discovered by surprise.


Note

Medium Risk
Benchmark harness changes combine an image bump, INT8-quantized all-reduce (not bit-identical to the prior path), and scheduling/graph sizing that can shift AgentX throughput and eval behavior.

Overview
Updates the qwen3.5-fp4-mi355x-sglang-agentic-mtp arm to SGLang ROCm v0.5.18 and retunes the MI355X AgentX launch script so it matches sibling recipes and the published MXFP4 cookbook.

In qwen3.5_fp4_mi355x_sglang_mtp.sh, multi-GPU collectives now use INT8 ROCm quick all-reduce (ROCM_QUICK_REDUCE_QUANTIZATION=INT8), EAGLE draft-extend goes through AITER unified attention (SGLANG_AITER_UNIFIED_DRAFT_EXTEND=1), prefill caps drop from 32768 to 16384 for --max-prefill-tokens and --chunked-prefill-size (aligned with the B200 script), and decode CUDA graphs are captured to min(2×CONC, 128) instead of min(CONC, 64) so they cover --max-running-requests at 2×CONC.

configs/amd-master.yaml pins the new image; perf-changelog.yaml records the sweep trigger.

Reviewed by Cursor Bugbot for commit ed29d46. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@github-actions

This comment was marked as outdated.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

yichiche and others added 2 commits August 26, 2026 16:44
…MXFP4 MI355X AgentX arm

Route the EAGLE draft-extend step through the AITER unified attention
kernel, and drop the CUDA graph sizing comment now that the changelog
entry carries the rationale.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…x-sglang-agentic-v0.5.18

# Conflicts:
#	perf-changelog.yaml
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant