Skip to content

fix(trtllm): add numa_cpu_bind to bind worker ranks per NUMA domain - #332

Open
Tabrizian wants to merge 1 commit into
NVIDIA:mainfrom
Tabrizian:fix/trtllm-numa-cpu-bind
Open

Tabrizian wants to merge 1 commit into
NVIDIA:mainfrom
Tabrizian:fix/trtllm-numa-cpu-bind

Conversation

@Tabrizian

@Tabrizian Tabrizian commented Aug 22, 2026 •

Copy link
Copy Markdown
Member

TrtllmBackend.get_srun_config() hardcodes cpu_bind="verbose,none" (src/srtctl/backends/trtllm.py), which disables CPU binding for every TRT-LLM run. On hosts with many NUMA domains the worker ranks are then free to migrate, and they contend.

Measurement

b300 (AMD EPYC 9575F, 8 NUMA domains, 8 gen ranks/node), GLM-5.2 NVFP4 disaggregated decode. Configs byte-identical; binding is the only variable:

decode p50
unbound (current behaviour) 70.4 ms
--cpu-bind=rank_ldom 14.1 ms (n=826, p99 14.2)
aggregated, same hardware 14.0 ms
reference cluster (2 NUMA domains) 13.9 ms

5.0x, landing on the aggregated figure.

Root cause

The cost is host-side input prep — prepare_for_spec_decode (dsa.py:1480-1487), which nsys shows is 99.8% pure CPU Python (0.4 ms of CUDA runtime in a 256 ms range). A standalone reproducer of those exact statements:

1 rank on an idle node 0.22 ms
8 ranks on one node, unbound 35–56 ms, skewed (two ranks stay fast)
8 ranks, bound 0.21 ms, uniform

The skew is what presents in production as a rotating straggler: all ranks stall together in the per-iteration MPI collectives waiting on whichever rank is slowest that iteration.

The comparison that isolates it: a 2-NUMA-domain reference host runs the same bench at 1.14–1.40 ms uniform despite being 2.2x slower single-threaded (31.2 µs vs 14.1 µs on a fixed CPU loop). It is NUMA topology, not CPU speed — and the b300 host is the faster machine once ranks stop migrating.

NIXL, MPI/UCX transport (15 µs synchronized allgather), GIL contention, request starvation and the disagg path itself were each excluded on measurement.

Change

Adds numa_cpu_bind: bool | None = None, mirroring the existing numa_memory_bind option:

  • None (default) — unchanged behaviour, no existing cluster is affected
  • True — --cpu-bind=verbose,rank_ldom (task i → NUMA domain i)
  • False — forces off

Notes for reviewers

  • rank_ldom requires the whole node; srun rejects it with Entire node must be allocated otherwise. Pair with use_exclusive_sbatch_directive (schema default is False).
  • I also tried wrapping the worker in numactl --cpunodebind — do not: it breaks MPI startup (7 of 8 mgmn_worker_node processes die and rank 0 hangs at MpiCommSession). Binding must come from srun.
  • --mem-bind=local is deliberately not included: it OOM-killed the prefill worker, whose host KV pool exceeds one NUMA domain, and a microbenchmark showed memory binding alone gives no benefit. First-touch already places pages locally once threads are CPU-bound.

TrtllmBackend.get_srun_config() hardcodes cpu_bind="verbose,none", which
disables CPU binding for every TRT-LLM run. On hosts with many NUMA domains the
worker ranks then migrate and contend.

Measured on b300 (AMD EPYC 9575F, 8 NUMA domains, 8 gen ranks/node) with GLM-5.2
NVFP4 disaggregated decode, configs byte-identical and binding the only variable:

    unbound                       decode p50 70.4 ms
    --cpu-bind=rank_ldom          decode p50 14.1 ms   (n=826, p99 14.2)
    aggregated, same hardware     decode p50 14.0 ms

5.0x, landing on the aggregated figure. The cost is host-side input prep
(prepare_for_spec_decode, dsa.py:1480-1487) -- 99.8% CPU Python, 0.22 ms with one
rank and 35-56 ms with eight unpinned ranks on one node -- and the unbound tail is
skewed, which presents as a rotating straggler.

Opt-in: numa_cpu_bind defaults to None, preserving today's behavior. Requires the
whole node (srun rejects rank_ldom otherwise), so pair it with
use_exclusive_sbatch_directive.
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
⚠️ Please upload report for BASE (main@c88ef73). Learn more about missing BASE report.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #332   +/-   ##
=======================================
  Coverage        ?   72.25%           
=======================================
  Files           ?       87           
  Lines           ?    12715           
  Branches        ?        0           
=======================================
  Hits            ?     9187           
  Misses          ?     3528           
  Partials        ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

# Requires the whole node to be allocated -- srun rejects rank_ldom with
# "Entire node must be allocated" otherwise -- so enable it together with
# use_exclusive_sbatch_directive.
numa_cpu_bind: bool | None = None

@richardhuo-nv richardhuo-nv Aug 22, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe set this as "srun_cpu_bind_method", and directly write the bind method, default is none?

srun_cpu_bind_method: rank_ldom

numa_cpu_bind sounds like we are going to use numactl to do something, given there is already a numa_memory_bind.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants