Skip to content

EP AllToAll: size dispatch on the sequences a rank owns under attention-DP - #175

Open
ajassani wants to merge 3 commits into
AMD-AGI:mainfrom
ajassani:aj/fix-ep-a2a-attention-dp
Open

ajassani wants to merge 3 commits into
AMD-AGI:mainfrom
ajassani:aj/fix-ep-a2a-attention-dp

Conversation

@ajassani

@ajassani ajassani commented Sep 22, 2026 •

Copy link
Copy Markdown

Summary

  • Attention-DP splits running requests across ranks. Expert dispatch was still sized on the replica-wide batch, so DP=8 at batch 512 priced an 8x AllToAll on every rank.
  • Size the payload with the existing attn_dp on the collective: ceil(batch / attn_dp). DP=1 is unchanged.

Test plan

  • pytest tests/unit/projection/test_expert_parallel_comm.py (6 passed)
  • Replica batch 512 + DP=8 should match batch 64 + DP=1, and stay well below unsplit batch 512

ajassani and others added 2 commits September 22, 2026 16:57
Attention-DP splits requests across ranks, not heads. Dispatch priced on the replica-wide batch made DP=8 send an 8x payload on every rank.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Adeem Jassani <ajassani@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Adeem Jassani <ajassani@amd.com>
Signed-off-by: Adeem Jassani <ajassani@amd.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants