Repository navigation
Price pipeline-parallel hops, the token return, and cross-node stage boundaries - #179
Open
araina-amd wants to merge 30 commits into
Open
araina-amd wants to merge 30 commits into
araina-amd wants to merge 30 commits into
Conversation
A measured anchor was being read past its own coverage, so a calibrated mode could come out less accurate than the analytical one -- the opposite of what "calibrated" implies. A prefill probe too short to fit a context curve is now refused, decode staying calibrated while prefill and TTFT fall back to simulation. A decode sweep that does not reach the batch being asked about is warned instead, there being no second source for it. Neither changes a projection; they make the anchor's coverage legible. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The warning claimed the extrapolation under-costs the step and over-reads throughput. That holds on the closed-form path and reverses under trace replay, so the sign belongs to the scheduling path rather than to the anchor. Say only what the anchor settles: past its widest rung the step cost is modelled, not measured. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The stack was priced from one top-k scale floored at a constant, and the floor is the wrong shape for the work it stands in for. A compressed layer does not attend over its context: it attends over a local window plus a pool holding one entry per compressed span. So the term that keeps growing with context is the indexer scoring the pool, not the attention, which stops growing once the pool passes the top-k. The schedule and its sizing geometry were already in the yaml and nothing read either. Prefill only; models declaring no schedule keep the floor. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The loader reads ``extends``; these files said ``bases``, which it does not look at. So each of them silently inherited nothing, leaving every field the base was there to supply at its dataclass default rather than at the value the file names. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two corrections to what a rank actually holds per token, both of which move the concurrency ceiling rather than any latency. Native sparse attention has to decide what to attend to before it can attend, and it scores one index key per past token. Those keys are a per-token cache in their own right sitting beside the KV cache, and they were not counted. The key is head-shared the way an MLA latent is, so tensor parallelism does not divide it, and it keeps the stack's dtype rather than ``kv_cache_dtype`` -- the index keys do not shrink when the KV cache is quantized. Decode context parallelism slices the tokens across ranks, and is the only axis that shrinks an MLA cache at all: tensor parallelism cannot, because the latent is shared across heads so every rank keeps a whole copy. Unlike attention-DP it shrinks one request's footprint, so it raises the ceiling even at the smallest batch. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The admission bound asks how many requests the KV pool can hold at once, and both of its inputs were being inferred when the deployment already reports them. ``--kv-pool-tokens`` takes the pool the engine actually allocated, as vLLM and SGLang print it at startup, so a memory-model error no longer shows up as an admission error. ``--workload-resident-tokens`` takes what one request holds while resident, having defaulted to the configured context, which is exact only for a fixed-length workpoint. A spread of lengths is length-biased, because a request occupies the pool for a time proportional to its length, so the length-biased mean is wanted rather than the arithmetic one. It is deliberately not discounted by the prefix-cache hit rate: a hit skips prefill compute, it does not free the blocks. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The two stations keep separate clocks, and only one of them was being fed. A closed-loop client retires on the decode clock and submits its replacement stamped with that time, but arrivals were released against the prefill clock. Whenever the prefill pool drained its queue its clock stopped where its last batch left it while decode's kept moving, so every request reissued during that window was invisible to the pool that has to prefill it. The population in flight collapsed, and with nothing to batch with the decode pool ran at minimum batch however many clients were offered. Flat is what made it survive review -- a split that cannot use concurrency looks like a finding about disaggregation rather than a stalled loop. Arrivals now come off the simulation clock, which is the one handoffs were already delivered on, and an idle prefill pool advances to meet its next request. Nothing is prefilled before its client asked for it. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The parameter split calls the engine regime-defining and always has, but REGIME_AXES did not list it, so every engine's anchor for a checkpoint hashed to one signature. regime_distance then called anchors from different engines a perfect match, and the bench cache, which keys on the same axes, could return one engine's measurement for another engine's run. Re-indexing could not separate them either, because the store preferred the signature recorded in meta -- hashed by whichever build harvested the artifact, over whichever axes that build had. So the signature is now recomputed and the recorded one kept as provenance. Unknown stays skippable, since every artifact in the store predates the axis. The case this is for is mori-sglang, which runs sglang's launch script and is not sglang: the two report very different KV pools because one shards the MLA latent across ranks where the other replicates it. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
… reuse The reuse store evicted by pure recency. That is not a simpler approximation of what a radix prefix cache does, it is the LRU pathology: blocks are touched in prefix order, so a corpus whose working set exceeds the pool is a cyclic sweep, and a cyclic sweep evicts every block exactly before it is reused. Engines do not do this -- a block cannot be dropped while a resident block continues from it, so the shared head survives pressure and the divergent tails are what go. Eviction is now leaf-first, least recently used among the leaves, which is the policy SGLang's radix cache implements. It does not rescue curves whose working set genuinely does not fit, which is a separate question. Separately the disaggregated branch warmed a prefix cache and discarded the summary, so every split row reported no hit rate at all -- reading as a split that reuses nothing rather than one whose reuse was never recorded, on exactly the rows the reuse question was being asked about. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A replayed ladder reported strong reuse up to the point the offered contexts stopped fitting the pool and none at all past it, and throughput collapsed with it. The hardware degrades there; the replay fell off a cliff. Two things were wrong, both in the order prefixes are resolved in. The pre-pass walked the trace oldest-first, which is the order an engine admits in only while everything offered fits. Past that point oldest-first walks the clients round-robin, so a client's context is always evicted before its next turn and every request pays a full reprefill -- the cliff is the ordering, not the cache and not the working set. A scheduler with a deep waiting queue picks off it by longest resident prefix instead, which is SGLang's default policy. That alone was not enough, because the freed slot was refilled from a global pointer. A closed loop holds one request per client, and serving one frees that client, whose next turn continues the context just served; handing the slot to another client turns the loop into a sliding window over arrival order and drops exactly the reuse it was about to collect. Below the pressure point the two orders touch the same blocks and the correction is invisible, which is what makes it a missing behaviour rather than a dial. Schedule order says nothing about when a request arrived, so each instance is still handed its stream in arrival order. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Most of what decided TTFT fidelity turned out to be about the measurement window rather than about physics, and none of it was visible in the cost model. TTFT from arrival, not from admission. The replay measured it from the moment the server admitted a request, on the grounds that a harness's clock starts after the request acquires a concurrency slot. Under a closed loop there are exactly as many clients as slots, so a request never waits on a client semaphore, and the time it then spends in the server's waiting queue is time the harness has already been counting. The opening transient is dropped by issue order. Every client fires at once, so the first requests queue against each other and carry the run's longest waits; because they waited, they are also the last to retire, so dropping a fraction of the earliest completions keeps every one of them. The harness excludes the same transient by advancing each lane and draining before it profiles. Only a real queue is reorderable. Admitting by longest resident prefix was handed the full client count as its window, which let it reorder a queue that was empty -- and that is the regime where reuse collapses, because the resident lanes evict each other. The window is now the backlog. And the run is bounded by a clock rather than a request count, which is how a fixed-concurrency harness is bounded. A budget split across lanes fixes how far each lane walks into its conversation, and an agentic turn's prompt grows with its position, so a budget-bounded replay offers a prompt-length trend with concurrency that the hardware never had. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two of the three mechanisms from the previous commit do not pay for themselves on the AgentX corpus, ablated one at a time over the model-engine pairs, so they are now off by default. Dropping the opening burst turned out to be wrong about the harness rather than wrong in the code. Its lane-advance requests are one-token requests: they move each lane to its live position and warm the cache, but they cost nothing and drain immediately, so when profiling starts every lane still fires a real request at once and that burst is in the reported average. The mechanism stays, because a harness that warms properly would need it, but this one does not. The resident cap is a genuinely better account of what a scheduler can reorder -- a policy cannot reorder an empty queue -- and it does buy the ITL pairs it was aimed at. It costs a throughput pair to get it, by pushing reuse at the low rungs further down than the hardware's went, so it sits behind --des-admit-backlog-only rather than in the default path. What remains from that commit is TTFT measured from the client's send instead of from server admission, which is not a tuning choice. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
``run_des`` accepted ``duration_ms`` and ``prefill_exclusive`` and handed neither to ``simulate_multi_instance``, which did not take them at all. Both are read in ``simulate_once``, so both worked on the fixed-sequence path and were silently inert on every run with a block cache to model -- a synthetic prefix pool or a trace to replay, which is every agentic replay we score. A parameter dropped one frame up is indistinguishable from one never passed. The clock meant every replay reported over its whole trace while the flag meant to bound it sat unused, so a ladder ran windows of varying length against a harness that runs the same window at every rung. The prefill policy meant SGLang and Atom replays scheduled vLLM's unified batch, which dissolves the herd a closed-loop population forms and ranks those engines' configurations backwards on TTFT, not merely imprecisely. Both regression tests assert at the ``run_des`` entry point the harness calls rather than the loop that implements the behaviour, since that is the frame where the parameter was lost. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The replay's admission gate reserved every request its whole context. A prefix-cache hit means those leading blocks are already in the pool -- some earlier request put them there and this one attends to the same physical KV -- so charging them again per request double-counts the one copy the engine keeps. Prefill already read the hit and skipped the cached suffix; only the reservation ignored it. At the block reuse an agentic corpus carries, that inflates a request's footprint by more than an order of magnitude, so the replay queued against a pool that was nowhere near full and put a whole request latency in front of TTFT. The credit is not refcounted -- a block's reservation is released when the request that warmed it retires, though the block stays resident for reuse. Both directions approximate a refcounted pool, and undercharging shared blocks lands far nearer the measured occupancy. A cold run is untouched, since ``cached_prefix`` is zero without a cache to hit. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
AITER ships its MLA, unified-attention and DeepSeek-V4 backends as separate flag values. An engine records whichever one its command line named, while the projector records the family it models, so an anchor and a target naming the same kernels were counted as a regime mismatch. The consequence was silent and total: the axis refused the anchor, the launcher fell through to the analytical projection, and the run reported benchmark mode while calibrating nothing. No harvest could clear it, because the spelling the engine records is the one the engine was asked for. ix_recipe.ATTENTION_BACKEND has always folded these onto one family; fold the same way here. Per-axis rather than in _canon, since a table applied everywhere would let unrelated values collide. Unknown spellings pass through unchanged. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A served harvest asked for MTP and measured a machine without it, through independent faults none of which had a visible symptom. The speculative flags were accepted on the command line and folded into the artifact's regime signature, but no branch ever passed them to the engine, so the server started with no draft head. The resulting anchor described a non-speculative machine while claiming to describe a speculative one, which nothing downstream can detect. The flags are not translations of each other and are now written per engine. The artifact then recorded no speculative keys at all, because AnchorStore discards the computed regime signature and rebuilds the regime from meta. A harvest fixed only per the paragraph above would still have produced a refused anchor. Acceptance was never pinned. Anchors serve dummy weights, and acceptance is a property of the weights: random ones accept at chance, so an unpinned harvest measures the non-speculative decode rate and files it under a speculative label. --speculative-acceptance-length pins the committed golden acceptance. Also let ATOM be driven at all: it is an SGLang derivative shipping the same client module, but its image carries no vLLM, so requiring vLLM's client left every ATOM regime unmeasurable for want of a load generator rather than for want of an engine. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The closed-loop replay issued each turn the instant the last one returned, so a population of N conversations was N requests permanently in flight. Real agentic traffic is not that: the same lanes show a 14-17s median gap between a response completing and the next turn arriving, because something on the other end is reading the answer. Modelled as an always-ready client, the server is saturated at every concurrency the ladder names, so throughput reads as the engine's ceiling rather than as what the offered load asks for, and the ladder stops distinguishing its own rungs. Think time restores the distinction. The prefix cache and the KV pool were also each being handed the whole allocation, which they share. A replay could therefore admit a batch against a pool that a cache it had already sized was occupying, and neither side knew. The cache is now sized to what the running requests leave free. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Nothing in an anchor recorded which GPU produced its timings, and no regime axis compared one, so a measurement transported across parts unchallenged. In this matrix that is GLM-5.2: its only anchors are MXFP4 Quark builds that load on gfx950 alone, and they were pricing an MI325X deployment. The curve read 4.3x the measured throughput and a twentieth of the measured TTFT, the worst in the matrix, while reporting itself calibrated. Treat the part as regime-defining rather than transportable, the conservative of the two readings: transport is possible in principle, but only once an anchor says what it ran on. Anchors harvested before the axis existed are not retired wholesale -- the store stamps the part at build time for the pairs whose part can be argued for. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
DeepSeek-V4 is priced from its per-layer compression schedule. GLM-5.2 and MiniMax-M3 select the same way on every layer, so neither had a schedule to read and both fell through to a top-k floored at 0.15 -- on the agentic corpus, nine to twenty-seven times the selection it names, and nothing at all for the indexing it stands in for. Selection and indexing move in opposite directions, so one multiplier cannot be either: selection reads topk entries whatever the prompt length, so its share falls away like topk/context, while indexing scores the whole pool and grows with context as dense attention does. Priced apart, GLM-5.2 is 0.032 of dense at 131k falling to 0.021 at 365k. The missing input was the indexer's query-head count, which sizes no cache and so was never recorded. sparse_indexer_cost_scale prices the kernel rather than the arithmetic, because the same selection runs fused through aiter on gfx950 and unfused through Torch on gfx942. Decode keeps its calibrated dense charge and pays only the excess over a fused kernel. Every other model in the matrix was audited against its released config and needed nothing. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
ruff-format wraps a call that already fits, and the lint job is the whole PR. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
The lint job formats the whole tree. A call that fits the line stays on it; one that does not is split. That is the whole of this commit. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
araina-amd
requested review from
JohnQinAMD,
jiejingzhangamd,
limou102 and
xiaobochen-amd
as code owners
October 1, 2026 22:16
vLLM and SGLang resolve quantization_config.quant_method to a toolchain name
and the benchmark records it verbatim, so an MXFP4 anchor comes back as
quark. Quark emits fp8 and fp4 alike, so the name says nothing about the
dtype, and every GLM-MXFP4 anchor on MI355X was refused on
weight_dtype('mxfp4'->'quark') against the checkpoint it was measured on.
Toolchain names now read as an unknown weight dtype, which regime_distance
skips, rather than a wrong one, which forces a mismatch.
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Speculative methods reach each engine by name. Collapsing every method to MTP sent EAGLE3 and DSpark harvests looking for draft layers inside checkpoints that have none: a MiniMax MXFP4 checkpoint under vLLM failed to load its MTP layers, and the EAGLE3 draft it was handed was never passed on. ATOM's image carries neither vLLM nor SGLang, so its load generator is the vLLM legacy client ATOM vendors. Every client now draws a fresh seed: all three default to a fixed one, so each repeat of a probe resent the same prompts and the prefix cache served them, and the prefill probe priced only the uncached tail of each prompt instead of the whole prompt. --trust-remote-code given through --server-args now reaches the client too, which loads the same tokenizer (Kimi-K3). When the median ITL sits well above mean TPOT the stream is coalescing tokens, and the decode step is read from mean TPOT instead. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Under attention-DP a rank runs every head of its own sequences, so dense decode attention over the whole context turns compute-bound, and that compute is what selection or compression removes. Charged dense, DeepSeek-V4-Pro on ATOM with attention-DP projected an ITL several times what was measured, while the same load without attention-DP projected close to it. DPA decode is now costed like prefill: from the compression schedule, else the uniform selection, else the configured scale, plus the unfused indexer's excess. INFERASIM_DPA_DECODE_DENSE=1 restores the dense charge for ablation. Anchors are transported by the model's whole decode step at the width the target verifies at, not by a bare forward pass. A measured step carries its fixed costs, and a compute-only ratio grew them with batch too: carried up from a single-client anchor, DeepSeek-V4-Pro read a per-token latency far above measured at higher concurrency. Within that shape a resident sequence adds its compressed pools and window, not a dense read of its context. The fitted KV stream scales with the target's KV width; GLM-MXFP4 on SGLang projected fp8 and bf16 KV alike where they measured clearly different throughput. A single-context anchor carries its KV delta from the analytic step. A prefill chunk is simulated at the context already resident, not at its own length. An anchor's attention-DP size is read from its server args when the meta does not record it. Mixed steps take a prefill batch, so lockstep ranks each attend over their own chunk. The hybrid-attention TTFT test now runs at one client. Under load TTFT also holds the wait behind other prefills, which grows faster than the prompt for queueing reasons, so a sublinear prefill can read superlinear at high concurrency. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A closed-loop replay can now run sessions rather than one request per slot. Each request waits on the previous turn of its stream and on any subagents it joins, plus its own recorded think_ms from the Mooncake trace; a session takes a client slot when one frees and holds it until its whole tree drains, which is the shared-sampler client AIPerf's AgentX lanes run. Concurrency then counts sessions, so the server admits a session's subagents up to its own sequence limit; capping the batch at the concurrency queued a subagent behind its sibling's whole decode, which dominated tail TTFT. A turn whose prompt extends the request it waits on is not routed until that resolves, or both would prefill. The replay stops issuing at the horizon and rates over the window plus the drain, and reports total tok/s (prompt plus generated) beside output tok/s. --max-num-seqs caps running sequences as the engine does. --des-client-idle-cap-ms mirrors the harness's system-idle guard, which pulls client timers forward when nothing is in flight. INFERASIM_DES_LOCKSTEP prices a step at the combined batch of that many schedulers stepping together, as SGLang attention-DP ranks do at every MoE layer. --uncached-prompt-latency-us adds an opt-in TTFT-only host term per prompt token the cache did not serve, capped by --uncached-prompt-latency-max-tokens; both default to 0. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
…l front end Under attention-DP the requests one prefill step carries belong to different ranks, which attend over their own share in parallel; only the MoE sees the whole step's tokens. INFERASIM_DES_DPA_SPREAD=1 prices a mixed step as that many chunks, one per rank, up to the attention-DP size. INFERASIM_DES_SERIAL_FRONTEND_US adds an opt-in serial front end: one process tokenizes every prompt in send order, at the given microseconds per prompt token, before the scheduler sees it. TTFT is still timed from the send, so a prompt queued behind others pays their tokenize. Both default off. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
INFERASIM_DCP_DECODE=1 divides decode attention by the decode context parallel size, since each DCP rank reads its own slice of every sequence's cache. INFERASIM_DECODE_SPARSE_GROWTH=1 is for engines whose decode kernels read the compressed or top-k cache: the dense charge sets the one-sequence level, and each further resident sequence adds its pools, window and indexer pass rather than a dense read of its context. Both default off; on the AgentX set the DCP switch made GLM-5.2 worse and changed Kimi-K3 little, so it stays an ablation. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Node sizes differ, so the cap on a warmup's GPUs is a fixed count. A small fixed allocation also keeps a warmup from waiting on a large one. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
araina-amd
force-pushed
the
araina/pp-stage-and-multinode-cost
branch
from
October 8, 2026 17:58
b32f9b3 to
7e1f977
Compare
Formatting only, from the pinned ruff-format hook that CI runs on all files. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A PP-n forward makes n hops (n-1 activation boundaries plus the sampled tokens' return to the first stage), and each pays the next stage's host scheduling on top of the send/recv. Price that with pp_stage_overhead_us, fitted to SGLang on MI355X, send boundaries that cross a node over the NIC, and once micro-batches are in flight (INFERASIM_PP_INFLIGHT > 1) charge each stage its scheduler host time in series with its forward (pp_host_us plus pp_host_per_req_us per request). Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
araina-amd
force-pushed
the
araina/pp-stage-and-multinode-cost
branch
from
October 8, 2026 18:51
7e1f977 to
2e18078
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
A pipeline-parallel forward was priced as (pp-1) activation send/recvs. On
SGLang that transfer is a small part of what a hop costs: the next stage's
scheduling and synchronisation on the host dominate it, and the last stage's
sampled tokens have to return to the first before that micro-batch can be
scheduled again. So PP layouts read faster than they run, and PP over two nodes
was priced as if every boundary stayed on xGMI.
This prices a PP-n forward as n hops (the n-1 activation boundaries plus the
token return), each with a fixed stage overhead on top of the transfer, and
sends boundaries that cross a node over the NIC. Once several micro-batches
are in flight, it also charges each stage the scheduler host time it can no
longer hide.
Stacked on #174 (uses its DES cost kernel and lockstep batching).
Type of change
Changes
Every hop pays the next stage's host time
pp_p2p_mschargespp_stage_overhead_usper hop in addition to thesend/recv, and adds the sampled tokens' return from the last stage to the
first. The default is fitted to SGLang on MI355X at one client over PP 2/4/8
on Llama-3.1-8B, Qwen3-14B-FP8 and Qwen3.6-35B-A3B; the send/recv is a small
fraction of it.
Boundaries that cross a node go over the NIC
boundaries on the NIC, and the token return crosses it too; the rest are
priced on xGMI with single-node arguments.
pp_nodes()reads the span fromINFERASIM_PP_NODESwhen the placement isgiven, otherwise from how many nodes
tp * ppGPUs fill.INFERASIM_COLL_HWtakes a calibrated system profile (a JSON object or apath to one), e.g. a fitted xGMI mesh bandwidth.
A full pipeline cannot hide the scheduler
upstream a stage never waits on its neighbour, so the batch preparation it
hid in that wait at one client lands in series with every stage's forward.
INFERASIM_PP_INFLIGHT> 1, the DES charges each stagepp_host_us + pp_host_per_req_us × running requestsper decode and mixedstep. Both were read from SGLang's overlap-off minus overlap-on step on
MI355X and agree across the three models above.
INFERASIM_PP_SYNC_ALPHAfactor,which only existed to absorb the missing host time.
Behaviour change
pp_stage_overhead_usis on by default, so every PP>1 projection now costsmore per forward. The in-flight host term is opt-in
(
INFERASIM_PP_INFLIGHT). Settingpp_stage_overhead_us=0recovers thetransfer-only hop.
Testing
tests/unit/projection/pass, unchanged from Refuse an out-of-regime anchor, and price sparse attention as selection plus indexing #174.1/2/4/8, C=1/16/64, ISL 1k and 8k) for the three models above.
Checklist: