Repository navigation
Refuse an out-of-regime anchor, and price sparse attention as selection plus indexing - #174
Open
araina-amd wants to merge 29 commits into
Open
araina-amd wants to merge 29 commits into
araina-amd wants to merge 29 commits into
Conversation
araina-amd
requested review from
JohnQinAMD,
jiejingzhangamd,
limou102 and
xiaobochen-amd
as code owners
September 21, 2026 18:01
A measured anchor was being read past its own coverage, so a calibrated mode could come out less accurate than the analytical one -- the opposite of what "calibrated" implies. A prefill probe too short to fit a context curve is now refused, decode staying calibrated while prefill and TTFT fall back to simulation. A decode sweep that does not reach the batch being asked about is warned instead, there being no second source for it. Neither changes a projection; they make the anchor's coverage legible. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The warning claimed the extrapolation under-costs the step and over-reads throughput. That holds on the closed-form path and reverses under trace replay, so the sign belongs to the scheduling path rather than to the anchor. Say only what the anchor settles: past its widest rung the step cost is modelled, not measured. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The stack was priced from one top-k scale floored at a constant, and the floor is the wrong shape for the work it stands in for. A compressed layer does not attend over its context: it attends over a local window plus a pool holding one entry per compressed span. So the term that keeps growing with context is the indexer scoring the pool, not the attention, which stops growing once the pool passes the top-k. The schedule and its sizing geometry were already in the yaml and nothing read either. Prefill only; models declaring no schedule keep the floor. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The loader reads ``extends``; these files said ``bases``, which it does not look at. So each of them silently inherited nothing, leaving every field the base was there to supply at its dataclass default rather than at the value the file names. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two corrections to what a rank actually holds per token, both of which move the concurrency ceiling rather than any latency. Native sparse attention has to decide what to attend to before it can attend, and it scores one index key per past token. Those keys are a per-token cache in their own right sitting beside the KV cache, and they were not counted. The key is head-shared the way an MLA latent is, so tensor parallelism does not divide it, and it keeps the stack's dtype rather than ``kv_cache_dtype`` -- the index keys do not shrink when the KV cache is quantized. Decode context parallelism slices the tokens across ranks, and is the only axis that shrinks an MLA cache at all: tensor parallelism cannot, because the latent is shared across heads so every rank keeps a whole copy. Unlike attention-DP it shrinks one request's footprint, so it raises the ceiling even at the smallest batch. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The admission bound asks how many requests the KV pool can hold at once, and both of its inputs were being inferred when the deployment already reports them. ``--kv-pool-tokens`` takes the pool the engine actually allocated, as vLLM and SGLang print it at startup, so a memory-model error no longer shows up as an admission error. ``--workload-resident-tokens`` takes what one request holds while resident, having defaulted to the configured context, which is exact only for a fixed-length workpoint. A spread of lengths is length-biased, because a request occupies the pool for a time proportional to its length, so the length-biased mean is wanted rather than the arithmetic one. It is deliberately not discounted by the prefix-cache hit rate: a hit skips prefill compute, it does not free the blocks. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The two stations keep separate clocks, and only one of them was being fed. A closed-loop client retires on the decode clock and submits its replacement stamped with that time, but arrivals were released against the prefill clock. Whenever the prefill pool drained its queue its clock stopped where its last batch left it while decode's kept moving, so every request reissued during that window was invisible to the pool that has to prefill it. The population in flight collapsed, and with nothing to batch with the decode pool ran at minimum batch however many clients were offered. Flat is what made it survive review -- a split that cannot use concurrency looks like a finding about disaggregation rather than a stalled loop. Arrivals now come off the simulation clock, which is the one handoffs were already delivered on, and an idle prefill pool advances to meet its next request. Nothing is prefilled before its client asked for it. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The parameter split calls the engine regime-defining and always has, but REGIME_AXES did not list it, so every engine's anchor for a checkpoint hashed to one signature. regime_distance then called anchors from different engines a perfect match, and the bench cache, which keys on the same axes, could return one engine's measurement for another engine's run. Re-indexing could not separate them either, because the store preferred the signature recorded in meta -- hashed by whichever build harvested the artifact, over whichever axes that build had. So the signature is now recomputed and the recorded one kept as provenance. Unknown stays skippable, since every artifact in the store predates the axis. The case this is for is mori-sglang, which runs sglang's launch script and is not sglang: the two report very different KV pools because one shards the MLA latent across ranks where the other replicates it. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
… reuse The reuse store evicted by pure recency. That is not a simpler approximation of what a radix prefix cache does, it is the LRU pathology: blocks are touched in prefix order, so a corpus whose working set exceeds the pool is a cyclic sweep, and a cyclic sweep evicts every block exactly before it is reused. Engines do not do this -- a block cannot be dropped while a resident block continues from it, so the shared head survives pressure and the divergent tails are what go. Eviction is now leaf-first, least recently used among the leaves, which is the policy SGLang's radix cache implements. It does not rescue curves whose working set genuinely does not fit, which is a separate question. Separately the disaggregated branch warmed a prefix cache and discarded the summary, so every split row reported no hit rate at all -- reading as a split that reuses nothing rather than one whose reuse was never recorded, on exactly the rows the reuse question was being asked about. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A replayed ladder reported strong reuse up to the point the offered contexts stopped fitting the pool and none at all past it, and throughput collapsed with it. The hardware degrades there; the replay fell off a cliff. Two things were wrong, both in the order prefixes are resolved in. The pre-pass walked the trace oldest-first, which is the order an engine admits in only while everything offered fits. Past that point oldest-first walks the clients round-robin, so a client's context is always evicted before its next turn and every request pays a full reprefill -- the cliff is the ordering, not the cache and not the working set. A scheduler with a deep waiting queue picks off it by longest resident prefix instead, which is SGLang's default policy. That alone was not enough, because the freed slot was refilled from a global pointer. A closed loop holds one request per client, and serving one frees that client, whose next turn continues the context just served; handing the slot to another client turns the loop into a sliding window over arrival order and drops exactly the reuse it was about to collect. Below the pressure point the two orders touch the same blocks and the correction is invisible, which is what makes it a missing behaviour rather than a dial. Schedule order says nothing about when a request arrived, so each instance is still handed its stream in arrival order. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Most of what decided TTFT fidelity turned out to be about the measurement window rather than about physics, and none of it was visible in the cost model. TTFT from arrival, not from admission. The replay measured it from the moment the server admitted a request, on the grounds that a harness's clock starts after the request acquires a concurrency slot. Under a closed loop there are exactly as many clients as slots, so a request never waits on a client semaphore, and the time it then spends in the server's waiting queue is time the harness has already been counting. The opening transient is dropped by issue order. Every client fires at once, so the first requests queue against each other and carry the run's longest waits; because they waited, they are also the last to retire, so dropping a fraction of the earliest completions keeps every one of them. The harness excludes the same transient by advancing each lane and draining before it profiles. Only a real queue is reorderable. Admitting by longest resident prefix was handed the full client count as its window, which let it reorder a queue that was empty -- and that is the regime where reuse collapses, because the resident lanes evict each other. The window is now the backlog. And the run is bounded by a clock rather than a request count, which is how a fixed-concurrency harness is bounded. A budget split across lanes fixes how far each lane walks into its conversation, and an agentic turn's prompt grows with its position, so a budget-bounded replay offers a prompt-length trend with concurrency that the hardware never had. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two of the three mechanisms from the previous commit do not pay for themselves on the AgentX corpus, ablated one at a time over the model-engine pairs, so they are now off by default. Dropping the opening burst turned out to be wrong about the harness rather than wrong in the code. Its lane-advance requests are one-token requests: they move each lane to its live position and warm the cache, but they cost nothing and drain immediately, so when profiling starts every lane still fires a real request at once and that burst is in the reported average. The mechanism stays, because a harness that warms properly would need it, but this one does not. The resident cap is a genuinely better account of what a scheduler can reorder -- a policy cannot reorder an empty queue -- and it does buy the ITL pairs it was aimed at. It costs a throughput pair to get it, by pushing reuse at the low rungs further down than the hardware's went, so it sits behind --des-admit-backlog-only rather than in the default path. What remains from that commit is TTFT measured from the client's send instead of from server admission, which is not a tuning choice. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
``run_des`` accepted ``duration_ms`` and ``prefill_exclusive`` and handed neither to ``simulate_multi_instance``, which did not take them at all. Both are read in ``simulate_once``, so both worked on the fixed-sequence path and were silently inert on every run with a block cache to model -- a synthetic prefix pool or a trace to replay, which is every agentic replay we score. A parameter dropped one frame up is indistinguishable from one never passed. The clock meant every replay reported over its whole trace while the flag meant to bound it sat unused, so a ladder ran windows of varying length against a harness that runs the same window at every rung. The prefill policy meant SGLang and Atom replays scheduled vLLM's unified batch, which dissolves the herd a closed-loop population forms and ranks those engines' configurations backwards on TTFT, not merely imprecisely. Both regression tests assert at the ``run_des`` entry point the harness calls rather than the loop that implements the behaviour, since that is the frame where the parameter was lost. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The replay's admission gate reserved every request its whole context. A prefix-cache hit means those leading blocks are already in the pool -- some earlier request put them there and this one attends to the same physical KV -- so charging them again per request double-counts the one copy the engine keeps. Prefill already read the hit and skipped the cached suffix; only the reservation ignored it. At the block reuse an agentic corpus carries, that inflates a request's footprint by more than an order of magnitude, so the replay queued against a pool that was nowhere near full and put a whole request latency in front of TTFT. The credit is not refcounted -- a block's reservation is released when the request that warmed it retires, though the block stays resident for reuse. Both directions approximate a refcounted pool, and undercharging shared blocks lands far nearer the measured occupancy. A cold run is untouched, since ``cached_prefix`` is zero without a cache to hit. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
AITER ships its MLA, unified-attention and DeepSeek-V4 backends as separate flag values. An engine records whichever one its command line named, while the projector records the family it models, so an anchor and a target naming the same kernels were counted as a regime mismatch. The consequence was silent and total: the axis refused the anchor, the launcher fell through to the analytical projection, and the run reported benchmark mode while calibrating nothing. No harvest could clear it, because the spelling the engine records is the one the engine was asked for. ix_recipe.ATTENTION_BACKEND has always folded these onto one family; fold the same way here. Per-axis rather than in _canon, since a table applied everywhere would let unrelated values collide. Unknown spellings pass through unchanged. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A served harvest asked for MTP and measured a machine without it, through independent faults none of which had a visible symptom. The speculative flags were accepted on the command line and folded into the artifact's regime signature, but no branch ever passed them to the engine, so the server started with no draft head. The resulting anchor described a non-speculative machine while claiming to describe a speculative one, which nothing downstream can detect. The flags are not translations of each other and are now written per engine. The artifact then recorded no speculative keys at all, because AnchorStore discards the computed regime signature and rebuilds the regime from meta. A harvest fixed only per the paragraph above would still have produced a refused anchor. Acceptance was never pinned. Anchors serve dummy weights, and acceptance is a property of the weights: random ones accept at chance, so an unpinned harvest measures the non-speculative decode rate and files it under a speculative label. --speculative-acceptance-length pins the committed golden acceptance. Also let ATOM be driven at all: it is an SGLang derivative shipping the same client module, but its image carries no vLLM, so requiring vLLM's client left every ATOM regime unmeasurable for want of a load generator rather than for want of an engine. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The closed-loop replay issued each turn the instant the last one returned, so a population of N conversations was N requests permanently in flight. Real agentic traffic is not that: the same lanes show a 14-17s median gap between a response completing and the next turn arriving, because something on the other end is reading the answer. Modelled as an always-ready client, the server is saturated at every concurrency the ladder names, so throughput reads as the engine's ceiling rather than as what the offered load asks for, and the ladder stops distinguishing its own rungs. Think time restores the distinction. The prefix cache and the KV pool were also each being handed the whole allocation, which they share. A replay could therefore admit a batch against a pool that a cache it had already sized was occupying, and neither side knew. The cache is now sized to what the running requests leave free. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Nothing in an anchor recorded which GPU produced its timings, and no regime axis compared one, so a measurement transported across parts unchallenged. In this matrix that is GLM-5.2: its only anchors are MXFP4 Quark builds that load on gfx950 alone, and they were pricing an MI325X deployment. The curve read 4.3x the measured throughput and a twentieth of the measured TTFT, the worst in the matrix, while reporting itself calibrated. Treat the part as regime-defining rather than transportable, the conservative of the two readings: transport is possible in principle, but only once an anchor says what it ran on. Anchors harvested before the axis existed are not retired wholesale -- the store stamps the part at build time for the pairs whose part can be argued for. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
DeepSeek-V4 is priced from its per-layer compression schedule. GLM-5.2 and MiniMax-M3 select the same way on every layer, so neither had a schedule to read and both fell through to a top-k floored at 0.15 -- on the agentic corpus, nine to twenty-seven times the selection it names, and nothing at all for the indexing it stands in for. Selection and indexing move in opposite directions, so one multiplier cannot be either: selection reads topk entries whatever the prompt length, so its share falls away like topk/context, while indexing scores the whole pool and grows with context as dense attention does. Priced apart, GLM-5.2 is 0.032 of dense at 131k falling to 0.021 at 365k. The missing input was the indexer's query-head count, which sizes no cache and so was never recorded. sparse_indexer_cost_scale prices the kernel rather than the arithmetic, because the same selection runs fused through aiter on gfx950 and unfused through Torch on gfx942. Decode keeps its calibrated dense charge and pays only the excess over a fused kernel. Every other model in the matrix was audited against its released config and needed nothing. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
ruff-format wraps a call that already fits, and the lint job is the whole PR. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
araina-amd
force-pushed
the
araina/anchor-regime-and-v4-attention
branch
from
September 21, 2026 18:08
bab0f1d to
0f8aa3e
Compare
The lint job formats the whole tree. A call that fits the line stays on it; one that does not is split. That is the whole of this commit. Signed-off-by: Anshu Raina <Anshu.Raina@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
6 of 8 tasks
vLLM and SGLang resolve quantization_config.quant_method to a toolchain name
and the benchmark records it verbatim, so an MXFP4 anchor comes back as
quark. Quark emits fp8 and fp4 alike, so the name says nothing about the
dtype, and every GLM-MXFP4 anchor on MI355X was refused on
weight_dtype('mxfp4'->'quark') against the checkpoint it was measured on.
Toolchain names now read as an unknown weight dtype, which regime_distance
skips, rather than a wrong one, which forces a mismatch.
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Speculative methods reach each engine by name. Collapsing every method to MTP sent EAGLE3 and DSpark harvests looking for draft layers inside checkpoints that have none: a MiniMax MXFP4 checkpoint under vLLM failed to load its MTP layers, and the EAGLE3 draft it was handed was never passed on. ATOM's image carries neither vLLM nor SGLang, so its load generator is the vLLM legacy client ATOM vendors. Every client now draws a fresh seed: all three default to a fixed one, so each repeat of a probe resent the same prompts and the prefix cache served them, and the prefill probe priced only the uncached tail of each prompt instead of the whole prompt. --trust-remote-code given through --server-args now reaches the client too, which loads the same tokenizer (Kimi-K3). When the median ITL sits well above mean TPOT the stream is coalescing tokens, and the decode step is read from mean TPOT instead. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Under attention-DP a rank runs every head of its own sequences, so dense decode attention over the whole context turns compute-bound, and that compute is what selection or compression removes. Charged dense, DeepSeek-V4-Pro on ATOM with attention-DP projected an ITL several times what was measured, while the same load without attention-DP projected close to it. DPA decode is now costed like prefill: from the compression schedule, else the uniform selection, else the configured scale, plus the unfused indexer's excess. INFERASIM_DPA_DECODE_DENSE=1 restores the dense charge for ablation. Anchors are transported by the model's whole decode step at the width the target verifies at, not by a bare forward pass. A measured step carries its fixed costs, and a compute-only ratio grew them with batch too: carried up from a single-client anchor, DeepSeek-V4-Pro read a per-token latency far above measured at higher concurrency. Within that shape a resident sequence adds its compressed pools and window, not a dense read of its context. The fitted KV stream scales with the target's KV width; GLM-MXFP4 on SGLang projected fp8 and bf16 KV alike where they measured clearly different throughput. A single-context anchor carries its KV delta from the analytic step. A prefill chunk is simulated at the context already resident, not at its own length. An anchor's attention-DP size is read from its server args when the meta does not record it. Mixed steps take a prefill batch, so lockstep ranks each attend over their own chunk. The hybrid-attention TTFT test now runs at one client. Under load TTFT also holds the wait behind other prefills, which grows faster than the prompt for queueing reasons, so a sublinear prefill can read superlinear at high concurrency. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A closed-loop replay can now run sessions rather than one request per slot. Each request waits on the previous turn of its stream and on any subagents it joins, plus its own recorded think_ms from the Mooncake trace; a session takes a client slot when one frees and holds it until its whole tree drains, which is the shared-sampler client AIPerf's AgentX lanes run. Concurrency then counts sessions, so the server admits a session's subagents up to its own sequence limit; capping the batch at the concurrency queued a subagent behind its sibling's whole decode, which dominated tail TTFT. A turn whose prompt extends the request it waits on is not routed until that resolves, or both would prefill. The replay stops issuing at the horizon and rates over the window plus the drain, and reports total tok/s (prompt plus generated) beside output tok/s. --max-num-seqs caps running sequences as the engine does. --des-client-idle-cap-ms mirrors the harness's system-idle guard, which pulls client timers forward when nothing is in flight. INFERASIM_DES_LOCKSTEP prices a step at the combined batch of that many schedulers stepping together, as SGLang attention-DP ranks do at every MoE layer. --uncached-prompt-latency-us adds an opt-in TTFT-only host term per prompt token the cache did not serve, capped by --uncached-prompt-latency-max-tokens; both default to 0. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
…l front end Under attention-DP the requests one prefill step carries belong to different ranks, which attend over their own share in parallel; only the MoE sees the whole step's tokens. INFERASIM_DES_DPA_SPREAD=1 prices a mixed step as that many chunks, one per rank, up to the attention-DP size. INFERASIM_DES_SERIAL_FRONTEND_US adds an opt-in serial front end: one process tokenizes every prompt in send order, at the given microseconds per prompt token, before the scheduler sees it. TTFT is still timed from the send, so a prompt queued behind others pays their tokenize. Both default off. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
INFERASIM_DCP_DECODE=1 divides decode attention by the decode context parallel size, since each DCP rank reads its own slice of every sequence's cache. INFERASIM_DECODE_SPARSE_GROWTH=1 is for engines whose decode kernels read the compressed or top-k cache: the dense charge sets the one-sequence level, and each further resident sequence adds its pools, window and indexer pass rather than a dense read of its context. Both default off; on the AgentX set the DCP switch made GLM-5.2 worse and changed Kimi-K3 little, so it stays an ablation. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Node sizes differ, so the cap on a warmup's GPUs is a fixed count. A small fixed allocation also keeps a warmup from waiting on a large one. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
araina-amd
force-pushed
the
araina/anchor-regime-and-v4-attention
branch
from
October 8, 2026 17:58
58da1d7 to
813e119
Compare
Formatting only, from the pinned ruff-format hook that CI runs on all files. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
An anchor is only a calibration if the measurement describes the deployment it is applied to. This branch makes that true for the serving engine, the GPU part, the attention-kernel family, the quantizer, and speculative decode; then it prices DeepSeek-V4, GLM-5.2, and MiniMax-M3 from the attention they actually run, including under attention-DP, instead of a top-k floor or a dense charge that is the wrong shape for the work.
The rest of the branch is the closed-loop replay that those projections are scored against. A conversation that never pauses, a prefix charged once per sharer, a cache that evicts the head before the tail, a disaggregated prefill pool on the wrong clock, and an agentic session replayed as independent requests are not small biases: they are the difference between a finding about the topology and a stalled loop.
Type of change
Changes
Regime: do not call a measurement a match when it is not
mori-sglangandsglangcould share a measurement. The two report very different KV pools because one shards the MLA latent and the other replicates it.quantization_config.quant_methodas the toolchain, so an MXFP4 anchor comes back asquark. Quark emits fp8 and fp4 alike, and every GLM-5.2-MXFP4 anchor on MI355X was refused on a weight-dtype mismatch against the checkpoint it was measured on. Toolchain names now read as an unknown weight dtype, whichregime_distanceskips, rather than a wrong one.Harvest: measure the deployment it is labelled with
AnchorStorethen rebuilt the regime from meta and dropped the speculative keys; acceptance was unpinned on dummy weights, so the measured rate was the non-speculative one filed under a speculative label.--trust-remote-codegiven through--server-argsreaches the client too, which loads the same tokenizer (Kimi-K3). When the median ITL sits well above mean TPOT the stream is coalescing tokens, and the decode step is read from mean TPOT instead.Attention: selection and indexing are not one multiplier
topk/context; indexing grows with context as dense attention does. Priced apart, GLM-5.2 is 0.032 of dense at 131k falling to 0.021 at 365k.sparse_indexer_n_headswas the missing input (it sizes no cache, so it was never recorded).sparse_indexer_cost_scaleprices the kernel: fused aiter on gfx950, unfused Torch on gfx942.INFERASIM_DPA_DECODE_DENSE=1restores the dense charge for ablation.basesnow inherit. The loader only readextends, so every field the base was there to supply was a dataclass default.Replay: a client that thinks, a pool that is shared, a cache that evicts leaves
duration_msandprefill_exclusivewere accepted atrun_desand never handed down, so every cached/trace run ignored the clock and SGLang/ATOM replays scheduled vLLM's unified batch.--kv-pool-tokensand--workload-resident-tokenstake the pool the engine allocated and the length-biased resident mean, so a memory-model error no longer shows up as an admission error.Replay: AgentX sessions the way the harness drives them
think_msfrom the Mooncake trace. A session takes a client slot when one frees and holds it until its whole tree drains, which is the shared-sampler client AIPerf's AgentX lanes run.--max-num-seqscaps running sequences as the engine does.--des-client-idle-cap-msmirrors the harness's system-idle guard, which pulls client timers forward when nothing is in flight.INFERASIM_DES_LOCKSTEPprices a step at the combined batch of that many schedulers stepping together, as SGLang attention-DP ranks do at every MoE layer; mixed steps take a prefill batch so lockstep ranks each attend over their own chunk.--uncached-prompt-latency-usadds an opt-in TTFT-only host term per prompt token the cache did not serve, capped by--uncached-prompt-latency-max-tokens; both default to 0.Testing
New and extended tests under
tests/unit/projection/for uniform sparse attention, hybrid compressed attention, speculative harvest, serving-anchor harvest, closed-loop think time / cache–pool sharing / measurement window / prefix cache, disaggregated reporting and DES scheduling, and anchor reuse across regime axes. All 448 tests undertests/unit/projection/pass.The hybrid-attention TTFT test now runs at one client. Under load TTFT also holds the wait behind other prefills, which grows faster than the prompt for queueing reasons,