Skip to content

Refuse an out-of-regime anchor, and price sparse attention as selection plus indexing - #174

Open
araina-amd wants to merge 29 commits into
mainfrom
araina/anchor-regime-and-v4-attention
Open

araina-amd wants to merge 29 commits into
mainfrom
araina/anchor-regime-and-v4-attention

Conversation

@araina-amd

@araina-amd araina-amd commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Description

An anchor is only a calibration if the measurement describes the deployment it is applied to. This branch makes that true for the serving engine, the GPU part, the attention-kernel family, the quantizer, and speculative decode; then it prices DeepSeek-V4, GLM-5.2, and MiniMax-M3 from the attention they actually run, including under attention-DP, instead of a top-k floor or a dense charge that is the wrong shape for the work.

The rest of the branch is the closed-loop replay that those projections are scored against. A conversation that never pauses, a prefix charged once per sharer, a cache that evicts the head before the tail, a disaggregated prefill pool on the wrong clock, and an agentic session replayed as independent requests are not small biases: they are the difference between a finding about the topology and a stalled loop.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)

Changes

Regime: do not call a measurement a match when it is not

  • Hardware part is regime-defining. Nothing in an anchor recorded which GPU produced it, so GLM-5.2's MXFP4 Quark timings (gfx950) were pricing an MI325X deployment. The curve read 4.3× the measured throughput and a twentieth of the measured TTFT, the worst in the matrix, while reporting itself calibrated. Unrecorded parts stay unknown rather than wrong; the store stamps the part at build time for the pairs whose part can be argued for.
  • Serving engine is a regime axis. Every engine's anchor for a checkpoint hashed to one signature, so mori-sglang and sglang could share a measurement. The two report very different KV pools because one shards the MLA latent and the other replicates it.
  • Attention backends compare by kernel family, not engine spelling. AITER's MLA / unified / DeepSeek-V4 flags were treated as distinct from the family the projector models, so the axis refused a matching harvest and the run reported benchmark mode while calibrating nothing.
  • A quantizer's name is not a dtype. vLLM and SGLang record quantization_config.quant_method as the toolchain, so an MXFP4 anchor comes back as quark. Quark emits fp8 and fp4 alike, and every GLM-5.2-MXFP4 anchor on MI355X was refused on a weight-dtype mismatch against the checkpoint it was measured on. Toolchain names now read as an unknown weight dtype, which regime_distance skips, rather than a wrong one.
  • Coverage is legible. A prefill probe too short to fit a context curve is refused (decode stays calibrated; prefill and TTFT fall back). A decode sweep that does not reach the asked batch is warned, not signed: the old warning claimed under-cost / over-throughput, which holds on the closed-form path and reverses under trace replay.

Harvest: measure the deployment it is labelled with

  • A speculative harvest now actually speculates. Flags were accepted into the regime signature but never passed to the engine; AnchorStore then rebuilt the regime from meta and dropped the speculative keys; acceptance was unpinned on dummy weights, so the measured rate was the non-speculative one filed under a speculative label.
  • Speculative methods reach each engine by name. Collapsing every method to MTP sent EAGLE3 and DSpark harvests looking for draft layers inside checkpoints that have none, and the EAGLE3 draft was never passed on.
  • ATOM is driveable. Its image carries neither vLLM nor SGLang, so its load generator is the vLLM legacy client ATOM vendors.
  • Every client draws a fresh seed. All three default to a fixed one, so each repeat of a probe resent the same prompts, the prefix cache served them, and the prefill probe priced only the uncached tail instead of the whole prompt.
  • --trust-remote-code given through --server-args reaches the client too, which loads the same tokenizer (Kimi-K3). When the median ITL sits well above mean TPOT the stream is coalescing tokens, and the decode step is read from mean TPOT instead.

Attention: selection and indexing are not one multiplier

  • DeepSeek-V4 is priced from the per-layer compression schedule the checkpoint ships. A compressed layer attends over a local window plus a pool, so the term that keeps growing with context is the indexer, not the attention.
  • GLM-5.2 and MiniMax-M3 select the same way on every layer, so they had no schedule and fell through to a top-k floored at 0.15 — nine to twenty-seven times the selection it names on the agentic corpus, and nothing for the indexing it stands in for. Selection falls like topk/context; indexing grows with context as dense attention does. Priced apart, GLM-5.2 is 0.032 of dense at 131k falling to 0.021 at 365k. sparse_indexer_n_heads was the missing input (it sizes no cache, so it was never recorded). sparse_indexer_cost_scale prices the kernel: fused aiter on gfx950, unfused Torch on gfx942.
  • Attention-DP decode is priced from the attention it runs. Under attention-DP a rank runs every head of its own sequences, so dense decode attention over the whole context turns compute-bound, and that compute is what selection or compression removes. Charged dense, DeepSeek-V4-Pro on ATOM with attention-DP projected an ITL several times what was measured. DPA decode is now costed like prefill: from the compression schedule, else the uniform selection, else the configured scale, plus the unfused indexer's excess. INFERASIM_DPA_DECODE_DENSE=1 restores the dense charge for ablation.
  • Anchors are transported by the model's whole decode step, at the width the target verifies at, not by a bare forward pass. A measured step carries its fixed costs, and a compute-only ratio grew them with batch too. Within that shape a resident sequence adds its compressed pools and window, not a dense read of its context.
  • The fitted KV stream scales with the target's KV width, so fp8 and bf16 KV no longer project alike. A single-context anchor carries its KV delta from the analytic step. A prefill chunk is simulated at the context already resident, not at its own length. An anchor's attention-DP size is read from its server args when the meta does not record it.
  • Configs that said bases now inherit. The loader only read extends, so every field the base was there to supply was a dataclass default.
  • Index keys sit beside the KV cache and are not divided by tensor parallelism; they keep the stack dtype when KV is quantized. Decode context parallelism is the axis that actually shrinks an MLA cache.

Replay: a client that thinks, a pool that is shared, a cache that evicts leaves

  • Think time. Closed-loop replay issued the next turn the instant the last one returned, so N conversations were N requests permanently in flight. The AgentX lanes show a 14–17s median gap. Without it the ladder is the engine's ceiling at every rung.
  • A shared prefix is reserved once. Prefill already skipped the hit; the admission gate charged every request its whole context, inflating the footprint by more than an order of magnitude and putting a whole request latency in front of TTFT.
  • The prefix cache and the KV pool share one allocation; each was being handed the whole of it.
  • Eviction is leaf-first (LRU among leaves), which is what a radix cache does. Recency-only eviction is the cyclic-sweep pathology: every block is dropped exactly before it is reused. A split now reports the reuse it already computed.
  • Admission is off the waiting queue by prefix, not by age. A closed loop holds one request per client; handing a freed slot to another client by arrival order drops the reuse it was about to collect.
  • The replay reports over the window the harness reports over. duration_ms and prefill_exclusive were accepted at run_des and never handed down, so every cached/trace run ignored the clock and SGLang/ATOM replays scheduled vLLM's unified batch.
  • Admission-window and warmup knobs are opt-in, on ablation against the AgentX corpus: the opening burst is the harness's one-token lane-advance, not a warm start, and still sits in the reported average.
  • A split's prefill pool is fed from the simulation clock. Arrivals stamped on the decode clock were released against a prefill clock that stopped when its queue drained, so the population in flight collapsed and decode ran at minimum batch — looking like a finding about disaggregation.
  • --kv-pool-tokens and --workload-resident-tokens take the pool the engine allocated and the length-biased resident mean, so a memory-model error no longer shows up as an admission error.

Replay: AgentX sessions the way the harness drives them

  • Sessions, not one request per slot. Each request waits on the previous turn of its stream and on any subagents it joins, plus its own recorded think_ms from the Mooncake trace. A session takes a client slot when one frees and holds it until its whole tree drains, which is the shared-sampler client AIPerf's AgentX lanes run.
  • Concurrency counts sessions, so the server admits a session's subagents up to its own sequence limit. Capping the batch at the concurrency queued a subagent behind its sibling's whole decode, which dominated tail TTFT.
  • A turn whose prompt extends the request it waits on is not routed until that resolves, or both would prefill.
  • The replay stops issuing at the horizon and rates over the window plus the drain, and reports total tok/s (prompt plus generated) beside output tok/s.
  • New knobs:
    • --max-num-seqs caps running sequences as the engine does.
    • --des-client-idle-cap-ms mirrors the harness's system-idle guard, which pulls client timers forward when nothing is in flight.
    • INFERASIM_DES_LOCKSTEP prices a step at the combined batch of that many schedulers stepping together, as SGLang attention-DP ranks do at every MoE layer; mixed steps take a prefill batch so lockstep ranks each attend over their own chunk.
    • --uncached-prompt-latency-us adds an opt-in TTFT-only host term per prompt token the cache did not serve, capped by --uncached-prompt-latency-max-tokens; both default to 0.

Testing

New and extended tests under tests/unit/projection/ for uniform sparse attention, hybrid compressed attention, speculative harvest, serving-anchor harvest, closed-loop think time / cache–pool sharing / measurement window / prefix cache, disaggregated reporting and DES scheduling, and anchor reuse across regime axes. All 448 tests under tests/unit/projection/ pass.

The hybrid-attention TTFT test now runs at one client. Under load TTFT also holds the wait behind other prefills, which grows faster than the prompt for queueing reasons,

araina-amd and others added 20 commits September 21, 2026 18:07
A measured anchor was being read past its own coverage, so a calibrated
mode could come out less accurate than the analytical one -- the opposite of
what "calibrated" implies.

A prefill probe too short to fit a context curve is now refused, decode
staying calibrated while prefill and TTFT fall back to simulation. A decode
sweep that does not reach the batch being asked about is warned instead, there
being no second source for it. Neither changes a projection; they make the
anchor's coverage legible.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The warning claimed the extrapolation under-costs the step and over-reads
throughput. That holds on the closed-form path and reverses under trace
replay, so the sign belongs to the scheduling path rather than to the anchor.
Say only what the anchor settles: past its widest rung the step cost is
modelled, not measured.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The stack was priced from one top-k scale floored at a constant, and the
floor is the wrong shape for the work it stands in for. A compressed layer
does not attend over its context: it attends over a local window plus a pool
holding one entry per compressed span. So the term that keeps growing with
context is the indexer scoring the pool, not the attention, which stops
growing once the pool passes the top-k.

The schedule and its sizing geometry were already in the yaml and nothing read
either. Prefill only; models declaring no schedule keep the floor.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The loader reads ``extends``; these files said ``bases``, which it does not
look at. So each of them silently inherited nothing, leaving every field the
base was there to supply at its dataclass default rather than at the value the
file names.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two corrections to what a rank actually holds per token, both of which move
the concurrency ceiling rather than any latency.

Native sparse attention has to decide what to attend to before it can attend,
and it scores one index key per past token. Those keys are a per-token cache
in their own right sitting beside the KV cache, and they were not counted. The
key is head-shared the way an MLA latent is, so tensor parallelism does not
divide it, and it keeps the stack's dtype rather than ``kv_cache_dtype`` --
the index keys do not shrink when the KV cache is quantized.

Decode context parallelism slices the tokens across ranks, and is the only
axis that shrinks an MLA cache at all: tensor parallelism cannot, because the
latent is shared across heads so every rank keeps a whole copy. Unlike
attention-DP it shrinks one request's footprint, so it raises the ceiling even
at the smallest batch.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The admission bound asks how many requests the KV pool can hold at once,
and both of its inputs were being inferred when the deployment already reports
them.

``--kv-pool-tokens`` takes the pool the engine actually allocated, as vLLM and
SGLang print it at startup, so a memory-model error no longer shows up as an
admission error.

``--workload-resident-tokens`` takes what one request holds while resident,
having defaulted to the configured context, which is exact only for a
fixed-length workpoint. A spread of lengths is length-biased, because a
request occupies the pool for a time proportional to its length, so the
length-biased mean is wanted rather than the arithmetic one. It is
deliberately not discounted by the prefix-cache hit rate: a hit skips prefill
compute, it does not free the blocks.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The two stations keep separate clocks, and only one of them was being fed.
A closed-loop client retires on the decode clock and submits its replacement
stamped with that time, but arrivals were released against the prefill clock.
Whenever the prefill pool drained its queue its clock stopped where its last
batch left it while decode's kept moving, so every request reissued during
that window was invisible to the pool that has to prefill it.

The population in flight collapsed, and with nothing to batch with the decode
pool ran at minimum batch however many clients were offered. Flat is what made
it survive review -- a split that cannot use concurrency looks like a finding
about disaggregation rather than a stalled loop.

Arrivals now come off the simulation clock, which is the one handoffs were
already delivered on, and an idle prefill pool advances to meet its next
request. Nothing is prefilled before its client asked for it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The parameter split calls the engine regime-defining and always has, but
REGIME_AXES did not list it, so every engine's anchor for a checkpoint hashed
to one signature. regime_distance then called anchors from different engines a
perfect match, and the bench cache, which keys on the same axes, could return
one engine's measurement for another engine's run.

Re-indexing could not separate them either, because the store preferred the
signature recorded in meta -- hashed by whichever build harvested the
artifact, over whichever axes that build had. So the signature is now
recomputed and the recorded one kept as provenance. Unknown stays skippable,
since every artifact in the store predates the axis.

The case this is for is mori-sglang, which runs sglang's launch script and is
not sglang: the two report very different KV pools because one shards the MLA
latent across ranks where the other replicates it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
… reuse

The reuse store evicted by pure recency. That is not a simpler
approximation of what a radix prefix cache does, it is the LRU pathology:
blocks are touched in prefix order, so a corpus whose working set exceeds the
pool is a cyclic sweep, and a cyclic sweep evicts every block exactly before
it is reused. Engines do not do this -- a block cannot be dropped while a
resident block continues from it, so the shared head survives pressure and the
divergent tails are what go. Eviction is now leaf-first, least recently used
among the leaves, which is the policy SGLang's radix cache implements. It does
not rescue curves whose working set genuinely does not fit, which is a
separate question.

Separately the disaggregated branch warmed a prefix cache and discarded the
summary, so every split row reported no hit rate at all -- reading as a split
that reuses nothing rather than one whose reuse was never recorded, on exactly
the rows the reuse question was being asked about.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A replayed ladder reported strong reuse up to the point the offered
contexts stopped fitting the pool and none at all past it, and throughput
collapsed with it. The hardware degrades there; the replay fell off a cliff.
Two things were wrong, both in the order prefixes are resolved in.

The pre-pass walked the trace oldest-first, which is the order an engine
admits in only while everything offered fits. Past that point oldest-first
walks the clients round-robin, so a client's context is always evicted before
its next turn and every request pays a full reprefill -- the cliff is the
ordering, not the cache and not the working set. A scheduler with a deep
waiting queue picks off it by longest resident prefix instead, which is
SGLang's default policy.

That alone was not enough, because the freed slot was refilled from a global
pointer. A closed loop holds one request per client, and serving one frees
that client, whose next turn continues the context just served; handing the
slot to another client turns the loop into a sliding window over arrival order
and drops exactly the reuse it was about to collect.

Below the pressure point the two orders touch the same blocks and the
correction is invisible, which is what makes it a missing behaviour rather
than a dial. Schedule order says nothing about when a request arrived, so each
instance is still handed its stream in arrival order.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Most of what decided TTFT fidelity turned out to be about the measurement
window rather than about physics, and none of it was visible in the cost
model.

TTFT from arrival, not from admission. The replay measured it from the moment
the server admitted a request, on the grounds that a harness's clock starts
after the request acquires a concurrency slot. Under a closed loop there are
exactly as many clients as slots, so a request never waits on a client
semaphore, and the time it then spends in the server's waiting queue is time
the harness has already been counting.

The opening transient is dropped by issue order. Every client fires at once,
so the first requests queue against each other and carry the run's longest
waits; because they waited, they are also the last to retire, so dropping a
fraction of the earliest completions keeps every one of them. The harness
excludes the same transient by advancing each lane and draining before it
profiles.

Only a real queue is reorderable. Admitting by longest resident prefix was
handed the full client count as its window, which let it reorder a queue that
was empty -- and that is the regime where reuse collapses, because the
resident lanes evict each other. The window is now the backlog.

And the run is bounded by a clock rather than a request count, which is how a
fixed-concurrency harness is bounded. A budget split across lanes fixes how
far each lane walks into its conversation, and an agentic turn's prompt grows
with its position, so a budget-bounded replay offers a prompt-length trend
with concurrency that the hardware never had.

Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Two of the three mechanisms from the previous commit do not pay for
themselves on the AgentX corpus, ablated one at a time over the model-engine
pairs, so they are now off by default.

Dropping the opening burst turned out to be wrong about the harness rather
than wrong in the code. Its lane-advance requests are one-token requests: they
move each lane to its live position and warm the cache, but they cost nothing
and drain immediately, so when profiling starts every lane still fires a real
request at once and that burst is in the reported average. The mechanism
stays, because a harness that warms properly would need it, but this one does
not.

The resident cap is a genuinely better account of what a scheduler can reorder
-- a policy cannot reorder an empty queue -- and it does buy the ITL pairs it
was aimed at. It costs a throughput pair to get it, by pushing reuse at the
low rungs further down than the hardware's went, so it sits behind
--des-admit-backlog-only rather than in the default path.

What remains from that commit is TTFT measured from the client's send instead
of from server admission, which is not a tuning choice.

Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
``run_des`` accepted ``duration_ms`` and ``prefill_exclusive`` and handed
neither to ``simulate_multi_instance``, which did not take them at all. Both
are read in ``simulate_once``, so both worked on the fixed-sequence path and
were silently inert on every run with a block cache to model -- a synthetic
prefix pool or a trace to replay, which is every agentic replay we score.

A parameter dropped one frame up is indistinguishable from one never passed.
The clock meant every replay reported over its whole trace while the flag
meant to bound it sat unused, so a ladder ran windows of varying length
against a harness that runs the same window at every rung. The prefill policy
meant SGLang and Atom replays scheduled vLLM's unified batch, which dissolves
the herd a closed-loop population forms and ranks those engines'
configurations backwards on TTFT, not merely imprecisely.

Both regression tests assert at the ``run_des`` entry point the harness calls
rather than the loop that implements the behaviour, since that is the frame
where the parameter was lost.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The replay's admission gate reserved every request its whole context. A
prefix-cache hit means those leading blocks are already in the pool -- some
earlier request put them there and this one attends to the same physical KV --
so charging them again per request double-counts the one copy the engine
keeps. Prefill already read the hit and skipped the cached suffix; only the
reservation ignored it.

At the block reuse an agentic corpus carries, that inflates a request's
footprint by more than an order of magnitude, so the replay queued against a
pool that was nowhere near full and put a whole request latency in front of
TTFT.

The credit is not refcounted -- a block's reservation is released when the
request that warmed it retires, though the block stays resident for reuse.
Both directions approximate a refcounted pool, and undercharging shared blocks
lands far nearer the measured occupancy. A cold run is untouched, since
``cached_prefix`` is zero without a cache to hit.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
AITER ships its MLA, unified-attention and DeepSeek-V4 backends as separate
flag values. An engine records whichever one its command line named, while the
projector records the family it models, so an anchor and a target naming the
same kernels were counted as a regime mismatch.

The consequence was silent and total: the axis refused the anchor, the
launcher fell through to the analytical projection, and the run reported
benchmark mode while calibrating nothing. No harvest could clear it, because
the spelling the engine records is the one the engine was asked for.

ix_recipe.ATTENTION_BACKEND has always folded these onto one family; fold the
same way here. Per-axis rather than in _canon, since a table applied
everywhere would let unrelated values collide. Unknown spellings pass through
unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A served harvest asked for MTP and measured a machine without it, through
independent faults none of which had a visible symptom.

The speculative flags were accepted on the command line and folded into the
artifact's regime signature, but no branch ever passed them to the engine, so
the server started with no draft head. The resulting anchor described a
non-speculative machine while claiming to describe a speculative one, which
nothing downstream can detect. The flags are not translations of each other
and are now written per engine.

The artifact then recorded no speculative keys at all, because AnchorStore
discards the computed regime signature and rebuilds the regime from meta. A
harvest fixed only per the paragraph above would still have produced a refused
anchor.

Acceptance was never pinned. Anchors serve dummy weights, and acceptance is a
property of the weights: random ones accept at chance, so an unpinned harvest
measures the non-speculative decode rate and files it under a speculative
label. --speculative-acceptance-length pins the committed golden acceptance.

Also let ATOM be driven at all: it is an SGLang derivative shipping the same
client module, but its image carries no vLLM, so requiring vLLM's client left
every ATOM regime unmeasurable for want of a load generator rather than for
want of an engine.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
The closed-loop replay issued each turn the instant the last one returned, so
a population of N conversations was N requests permanently in flight. Real
agentic traffic is not that: the same lanes show a 14-17s median gap between
a response completing and the next turn arriving, because something on the
other end is reading the answer.

Modelled as an always-ready client, the server is saturated at every
concurrency the ladder names, so throughput reads as the engine's ceiling
rather than as what the offered load asks for, and the ladder stops
distinguishing its own rungs. Think time restores the distinction.

The prefix cache and the KV pool were also each being handed the whole
allocation, which they share. A replay could therefore admit a batch against
a pool that a cache it had already sized was occupying, and neither side knew.
The cache is now sized to what the running requests leave free.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Nothing in an anchor recorded which GPU produced its timings, and no regime
axis compared one, so a measurement transported across parts unchallenged.
In this matrix that is GLM-5.2: its only anchors are MXFP4 Quark builds that
load on gfx950 alone, and they were pricing an MI325X deployment. The curve
read 4.3x the measured throughput and a twentieth of the measured TTFT, the
worst in the matrix, while reporting itself calibrated.

Treat the part as regime-defining rather than transportable, the conservative
of the two readings: transport is possible in principle, but only once an
anchor says what it ran on. Anchors harvested before the axis existed are not
retired wholesale -- the store stamps the part at build time for the pairs
whose part can be argued for.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
DeepSeek-V4 is priced from its per-layer compression schedule. GLM-5.2 and
MiniMax-M3 select the same way on every layer, so neither had a schedule to
read and both fell through to a top-k floored at 0.15 -- on the agentic
corpus, nine to twenty-seven times the selection it names, and nothing at all
for the indexing it stands in for.

Selection and indexing move in opposite directions, so one multiplier cannot
be either: selection reads topk entries whatever the prompt length, so its
share falls away like topk/context, while indexing scores the whole pool and
grows with context as dense attention does. Priced apart, GLM-5.2 is 0.032 of
dense at 131k falling to 0.021 at 365k. The missing input was the indexer's
query-head count, which sizes no cache and so was never recorded.

sparse_indexer_cost_scale prices the kernel rather than the arithmetic,
because the same selection runs fused through aiter on gfx950 and unfused
through Torch on gfx942. Decode keeps its calibrated dense charge and pays
only the excess over a fused kernel. Every other model in the matrix was
audited against its released config and needed nothing.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
ruff-format wraps a call that already fits, and the lint job is the whole PR.

Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@araina-amd
araina-amd force-pushed the araina/anchor-regime-and-v4-attention branch from bab0f1d to 0f8aa3e Compare September 21, 2026 18:08
The lint job formats the whole tree. A call that fits the line stays on
it; one that does not is split. That is the whole of this commit.

Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
araina-amd and others added 2 commits October 8, 2026 02:05
vLLM and SGLang resolve quantization_config.quant_method to a toolchain name
and the benchmark records it verbatim, so an MXFP4 anchor comes back as
quark. Quark emits fp8 and fp4 alike, so the name says nothing about the
dtype, and every GLM-MXFP4 anchor on MI355X was refused on
weight_dtype('mxfp4'->'quark') against the checkpoint it was measured on.

Toolchain names now read as an unknown weight dtype, which regime_distance
skips, rather than a wrong one, which forces a mismatch.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Speculative methods reach each engine by name. Collapsing every method to
MTP sent EAGLE3 and DSpark harvests looking for draft layers inside
checkpoints that have none: a MiniMax MXFP4 checkpoint under vLLM failed to
load its MTP layers, and the EAGLE3 draft it was handed was never passed on.

ATOM's image carries neither vLLM nor SGLang, so its load generator is the
vLLM legacy client ATOM vendors. Every client now draws a fresh seed: all
three default to a fixed one, so each repeat of a probe resent the same
prompts and the prefix cache served them, and the prefill probe priced only
the uncached tail of each prompt instead of the whole prompt.

--trust-remote-code given through --server-args now reaches the client too,
which loads the same tokenizer (Kimi-K3). When the median ITL sits well above
mean TPOT the stream is coalescing tokens, and the decode step is read from
mean TPOT instead.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
araina-amd and others added 5 commits October 8, 2026 02:05
Under attention-DP a rank runs every head of its own sequences, so dense
decode attention over the whole context turns compute-bound, and that compute
is what selection or compression removes. Charged dense, DeepSeek-V4-Pro on
ATOM with attention-DP projected an ITL several times what was measured,
while the same load without attention-DP projected close to it. DPA decode is
now costed like prefill: from the compression schedule, else the uniform
selection, else the configured scale, plus the unfused indexer's excess.
INFERASIM_DPA_DECODE_DENSE=1 restores the dense charge for ablation.

Anchors are transported by the model's whole decode step at the width the
target verifies at, not by a bare forward pass. A measured step carries its
fixed costs, and a compute-only ratio grew them with batch too: carried up
from a single-client anchor, DeepSeek-V4-Pro read a per-token latency far
above measured at higher concurrency. Within that shape a resident sequence
adds its compressed pools and window, not a dense read of its context.

The fitted KV stream scales with the target's KV width; GLM-MXFP4 on SGLang
projected fp8 and bf16 KV alike where they measured clearly different
throughput. A single-context anchor carries its KV delta from the analytic
step. A prefill chunk is simulated at the context already resident, not at
its own length. An anchor's attention-DP size is read from its server args
when the meta does not record it. Mixed steps take a prefill batch, so
lockstep ranks each attend over their own chunk.

The hybrid-attention TTFT test now runs at one client. Under load TTFT also
holds the wait behind other prefills, which grows faster than the prompt for
queueing reasons, so a sublinear prefill can read superlinear at high
concurrency.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
A closed-loop replay can now run sessions rather than one request per slot.
Each request waits on the previous turn of its stream and on any subagents
it joins, plus its own recorded think_ms from the Mooncake trace; a session
takes a client slot when one frees and holds it until its whole tree drains,
which is the shared-sampler client AIPerf's AgentX lanes run. Concurrency
then counts sessions, so the server admits a session's subagents up to its
own sequence limit; capping the batch at the concurrency queued a subagent
behind its sibling's whole decode, which dominated tail TTFT. A turn whose
prompt extends the request it waits on is not routed until that resolves,
or both would prefill. The replay stops issuing at the horizon and rates
over the window plus the drain, and reports total tok/s (prompt plus
generated) beside output tok/s.

--max-num-seqs caps running sequences as the engine does.
--des-client-idle-cap-ms mirrors the harness's system-idle guard, which
pulls client timers forward when nothing is in flight.
INFERASIM_DES_LOCKSTEP prices a step at the combined batch of that many
schedulers stepping together, as SGLang attention-DP ranks do at every MoE
layer. --uncached-prompt-latency-us adds an opt-in TTFT-only host term per
prompt token the cache did not serve, capped by
--uncached-prompt-latency-max-tokens; both default to 0.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
…l front end

Under attention-DP the requests one prefill step carries belong to different
ranks, which attend over their own share in parallel; only the MoE sees the
whole step's tokens. INFERASIM_DES_DPA_SPREAD=1 prices a mixed step as that
many chunks, one per rank, up to the attention-DP size.

INFERASIM_DES_SERIAL_FRONTEND_US adds an opt-in serial front end: one process
tokenizes every prompt in send order, at the given microseconds per prompt
token, before the scheduler sees it. TTFT is still timed from the send, so a
prompt queued behind others pays their tokenize. Both default off.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
INFERASIM_DCP_DECODE=1 divides decode attention by the decode context
parallel size, since each DCP rank reads its own slice of every sequence's
cache. INFERASIM_DECODE_SPARSE_GROWTH=1 is for engines whose decode kernels
read the compressed or top-k cache: the dense charge sets the one-sequence
level, and each further resident sequence adds its pools, window and indexer
pass rather than a dense read of its context. Both default off; on the AgentX
set the DCP switch made GLM-5.2 worse and changed Kimi-K3 little, so it stays
an ablation.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
Node sizes differ, so the cap on a warmup's GPUs is a fixed count. A small
fixed allocation also keeps a warmup from waiting on a large one.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>
@araina-amd
araina-amd force-pushed the araina/anchor-regime-and-v4-attention branch from 58da1d7 to 813e119 Compare October 8, 2026 17:58
Formatting only, from the pinned ruff-format hook that CI runs on all files.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Anshu Raina <Anshu.Raina@amd.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant