Conversation
TrtllmBackend.get_srun_config() hardcodes cpu_bind="verbose,none", which
disables CPU binding for every TRT-LLM run. On hosts with many NUMA domains the
worker ranks then migrate and contend.
Measured on b300 (AMD EPYC 9575F, 8 NUMA domains, 8 gen ranks/node) with GLM-5.2
NVFP4 disaggregated decode, configs byte-identical and binding the only variable:
unbound decode p50 70.4 ms
--cpu-bind=rank_ldom decode p50 14.1 ms (n=826, p99 14.2)
aggregated, same hardware decode p50 14.0 ms
5.0x, landing on the aggregated figure. The cost is host-side input prep
(prepare_for_spec_decode, dsa.py:1480-1487) -- 99.8% CPU Python, 0.22 ms with one
rank and 35-56 ms with eight unpinned ranks on one node -- and the unbound tail is
skewed, which presents as a rotating straggler.
Opt-in: numa_cpu_bind defaults to None, preserving today's behavior. Requires the
whole node (srun rejects rank_ldom otherwise), so pair it with
use_exclusive_sbatch_directive.
Tabrizian
requested review from
alec-flowers,
csahithi,
ishandhanani and
nlevin-ui
as code owners
August 22, 2026 06:56
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #332 +/- ##
=======================================
Coverage ? 72.25%
=======================================
Files ? 87
Lines ? 12715
Branches ? 0
=======================================
Hits ? 9187
Misses ? 3528
Partials ? 0 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
| # Requires the whole node to be allocated -- srun rejects rank_ldom with | ||
| # "Entire node must be allocated" otherwise -- so enable it together with | ||
| # use_exclusive_sbatch_directive. | ||
| numa_cpu_bind: bool | None = None |
Collaborator
There was a problem hiding this comment.
maybe set this as "srun_cpu_bind_method", and directly write the bind method, default is none?
srun_cpu_bind_method: rank_ldom
numa_cpu_bind sounds like we are going to use numactl to do something, given there is already a numa_memory_bind.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TrtllmBackend.get_srun_config()hardcodescpu_bind="verbose,none"(src/srtctl/backends/trtllm.py), which disables CPU binding for every TRT-LLM run. On hosts with many NUMA domains the worker ranks are then free to migrate, and they contend.Measurement
b300 (AMD EPYC 9575F, 8 NUMA domains, 8 gen ranks/node), GLM-5.2 NVFP4 disaggregated decode. Configs byte-identical; binding is the only variable:
--cpu-bind=rank_ldom5.0x, landing on the aggregated figure.
Root cause
The cost is host-side input prep —
prepare_for_spec_decode(dsa.py:1480-1487), which nsys shows is 99.8% pure CPU Python (0.4 ms of CUDA runtime in a 256 ms range). A standalone reproducer of those exact statements:The skew is what presents in production as a rotating straggler: all ranks stall together in the per-iteration MPI collectives waiting on whichever rank is slowest that iteration.
The comparison that isolates it: a 2-NUMA-domain reference host runs the same bench at 1.14–1.40 ms uniform despite being 2.2x slower single-threaded (31.2 µs vs 14.1 µs on a fixed CPU loop). It is NUMA topology, not CPU speed — and the b300 host is the faster machine once ranks stop migrating.
NIXL, MPI/UCX transport (15 µs synchronized allgather), GIL contention, request starvation and the disagg path itself were each excluded on measurement.
Change
Adds
numa_cpu_bind: bool | None = None, mirroring the existingnuma_memory_bindoption:None(default) — unchanged behaviour, no existing cluster is affectedTrue—--cpu-bind=verbose,rank_ldom(task i → NUMA domain i)False— forces offNotes for reviewers
rank_ldomrequires the whole node; srun rejects it withEntire node must be allocatedotherwise. Pair withuse_exclusive_sbatch_directive(schema default is False).numactl --cpunodebind— do not: it breaks MPI startup (7 of 8mgmn_worker_nodeprocesses die and rank 0 hangs atMpiCommSession). Binding must come from srun.--mem-bind=localis deliberately not included: it OOM-killed the prefill worker, whose host KV pool exceeds one NUMA domain, and a microbenchmark showed memory binding alone gives no benefit. First-touch already places pages locally once threads are CPU-bound.