-
Notifications
You must be signed in to change notification settings - Fork 273
Pull requests: SemiAnalysisAI/InferenceX
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[AMD][MI35X] Bump Qwen3.5 MXFP4 MI355X SGLang AgentX to v0.5.18 and retune all-reduce, prefill, and CUDA graph
full-sweep-fail-fast
#2737
opened Aug 26, 2026 by
yichiche
Collaborator
Loading…
[CI] Import the DCGM exporter via enroot registry syntax / 用 enroot registry 语法导入 DCGM exporter
#2735
opened Aug 26, 2026 by
edwingao28
Collaborator
Loading…
[Klaud Cold] minimaxm3-fp4-mi355x-atom-agentic-mtp: day-zero MiniMax-M3 MXFP4 ATOM AgentX recipe on MI355X / MI355X 上 MiniMax-M3 MXFP4 ATOM AgentX 首发配方
full-sweep-fail-fast
#2733
opened Aug 26, 2026 by
functionstackx
Collaborator
Loading…
Migrate MI300X runners to the Barite AMD cluster
#2732
opened Aug 25, 2026 by
cquil11
Collaborator
Loading…
test: validate upstream-native vLLM Router topologies
#2731
opened Aug 25, 2026 by
cquil11
Collaborator
Loading…
perf(gb300): Refresh Qwen3.5 FP4 GB300 dynamo-trt STP/MTP recipes from srt-slurm / 刷新基于 srt-slurm 的 Qwen3.5 FP4 GB300 dynamo-trt STP/MTP 配方
full-sweep-fail-fast
#2730
opened Aug 25, 2026 by
richardhuo-nv
Collaborator
Loading…
2 of 3 tasks
feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0821 (mirror of #2723)
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-enabled
#2725
opened Aug 25, 2026 by
seungrokj
Collaborator
Loading…
[WIP][AMD][AgentX] Add Qwen3.8 FP8 MI355X two-node vLLM
agentx-fast
Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable
sweep-enabled
#2724
opened Aug 25, 2026 by
haic0
Collaborator
Loading…
6 of 7 tasks
feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0821
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-enabled
#2723
opened Aug 25, 2026 by
zejunchen-zejun
Collaborator
Loading…
feat(dsv4): add B200 DeepSeek-V4-Pro dynamo-trt 8k1k disagg throughput recipes / 新增 B200 DeepSeek-V4-Pro dynamo-trt 8k1k 分离式吞吐量配置
full-sweep-fail-fast
#2721
opened Aug 24, 2026 by
richardhuo-nv
Collaborator
Loading…
3 of 4 tasks
feat(config): add GLM-5.2 GB300 AgentX concurrency-1 disaggregated point / 添加 GLM-5.2 GB300 AgentX 并发度 1 分离式配置
full-sweep-enabled
#2720
opened Aug 24, 2026 by
RohitNagraj
Collaborator
Loading…
4 of 9 tasks
Tune B200 DSV4 SGLang AgentX prefill scheduling / 调整 B200 DSV4 SGLang AgentX 预填充调度
full-sweep-fail-fast
#2718
opened Aug 24, 2026 by
nvpohanh
Collaborator
Loading…
fix: fold DP_SIZE into single-node num_gpus for per-GPU throughput
#2715
opened Aug 24, 2026 by
aistackdev
Loading…
config(dsv4): constrain OpenMP threads for GB300 AgentX / 限制 GB300 AgentX 的 OpenMP 线程数
full-sweep-enabled
#2711
opened Aug 24, 2026 by
RohitNagraj
Collaborator
Loading…
4 of 9 tasks
[AgentX] DeepSeek-V4 B300 SGLang update
full-sweep-enabled
#2704
opened Aug 21, 2026 by
yhyang201
Collaborator
Loading…
[AMD][MI35X] Add HiCache and TP2/EP1 arms to the Qwen3.5 MXFP4 MI355X AgentX sweep
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-enabled
#2693
opened Aug 20, 2026 by
yichiche
Collaborator
Loading…
[AMD] Refresh DeepSeek-R1 MI355X SGLang image / [AMD] 更新 DeepSeek-R1 MI355X SGLang 镜像
all-evals
Expand eval selection to every fixed-sequence config
full-sweep-fail-fast
#2691
opened Aug 20, 2026 by
Oseltamivir
Collaborator
Loading…
[Power] Require B200 and B300 multinode telemetry / 强制启用 B200 和 B300 多节点功耗采集
full-sweep-enabled
#2688
opened Aug 20, 2026 by
edwingao28
Collaborator
Loading…
Test AIPerf empty-content TTFT on MiniMax TRT
full-sweep-enabled
#2682
opened Aug 19, 2026 by
cquil11
Collaborator
Loading…
[DO NOT MERGE][Test] DSV4 FP4 B300 SGLang AgentX DEP8 c512 on nightly (HiCache + MegaMoE FP4-act)
sweep-enabled
#2680
opened Aug 19, 2026 by
yhyang201
Collaborator
Loading…
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.