Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
b0254de
Say when an anchor is being read outside the regime it was measured in
araina-amd Sep 16, 2026
62d49cf
Do not assert a direction for the decode batch extrapolation
araina-amd Sep 16, 2026
5de2511
Cost DeepSeek-V4 attention from the schedule the checkpoint ships
araina-amd Sep 16, 2026
e4134d6
Load the base config these models say they extend
araina-amd Sep 16, 2026
06ce624
Count the index keys a sparse model caches, and the tokens DCP splits
araina-amd Sep 16, 2026
27e00b6
Let a run be told the pool it is being validated against
araina-amd Sep 16, 2026
18f2ba7
Feed a split's prefill pool from the clock the whole run is on
araina-amd Sep 16, 2026
06ea1d6
Make the serving engine a regime axis
araina-amd Sep 16, 2026
0983c0d
Evict the replay's block cache leaf-first, and let a split report its…
araina-amd Sep 16, 2026
7a73b36
Admit off the replay's waiting queue by prefix, not by age
araina-amd Sep 16, 2026
586dd54
Report the replay over the window the harness reports over
araina-amd Sep 17, 2026
c43f856
Make the admission-window and warmup knobs opt-in, on the evidence
araina-amd Sep 17, 2026
ab1642c
Carry the clock and the prefill policy into the replay driver
araina-amd Sep 17, 2026
5d8680e
Charge a shared prefix to the pool once, not once per sharer
araina-amd Sep 17, 2026
46a127a
Compare attention backends by kernel family, not by engine spelling
araina-amd Sep 18, 2026
0c19edc
Harvest a speculative anchor that actually speculates
araina-amd Sep 18, 2026
cf46261
Replay a conversation as a client that thinks, not one that never stops
araina-amd Sep 20, 2026
d3899f5
Refuse an anchor measured on a different accelerator
araina-amd Sep 20, 2026
3082c6d
Charge a uniform top-k model for choosing, not just for reading
araina-amd Sep 20, 2026
0f8aa3e
Write the every-layer probe on one line
araina-amd Sep 21, 2026
bc28436
Let ruff wrap the projection files it already owns
araina-amd Sep 21, 2026
71969a6
Do not refuse an anchor over its quantizer's name
araina-amd Oct 1, 2026
ac1e110
Make a harvest measure the deployment it is labelled with
araina-amd Oct 1, 2026
69ccbb1
Price V4 and attention-DP decode from the attention they run
araina-amd Oct 1, 2026
7da65c8
Replay AgentX sessions the way the harness drives them
araina-amd Oct 1, 2026
6d49ce8
Spread an attention-DP prefill step over its ranks, and model a seria…
araina-amd Oct 6, 2026
6a95afe
Add opt-in DCP and sparse-growth decode attention charges
araina-amd Oct 6, 2026
813e119
State the warmup GPU cap as a count, not a share of a node
araina-amd Oct 6, 2026
fdfd15e
Let ruff format the projection files the replay and V4 work touched
araina-amd Oct 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 2 additions & 5 deletions infera/projection/agents/tuning_agent/inference_tuning.py
Original file line number Diff line number Diff line change
Expand Up @@ -635,9 +635,7 @@ def validate_inference(
if pool_dp < 1:
return False, f"{pool}_attention_dp={pool_dp} must be >= 1"
if pool_dp > 1 and pool_tp % pool_dp:
return False, (
f"{pool}_attention_dp={pool_dp} must divide {pool}_tp={pool_tp}"
)
return False, (f"{pool}_attention_dp={pool_dp} must divide {pool}_tp={pool_tp}")
for pool, pool_tp, pool_ep in (
("prefill", p_tp, cfg.prefill_ep),
("decode", d_tp, cfg.decode_ep),
Expand All @@ -648,8 +646,7 @@ def validate_inference(
return False, f"{pool}_ep={pool_ep} not in legal EP set {legality.ep}"
if pool_ep > 1 and (pool_tp * cfg.pp) % pool_ep:
return False, (
f"{pool}_ep={pool_ep} must divide the {pool} pool's "
f"{pool_tp * cfg.pp} ranks"
f"{pool}_ep={pool_ep} must divide the {pool} pool's {pool_tp * cfg.pp} ranks"
)
prefill_gpus = p_tp * cfg.pp * cfg.prefill_replicas
decode_gpus = d_tp * cfg.pp * cfg.decode_replicas
Expand Down
177 changes: 177 additions & 0 deletions infera/projection/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -528,6 +528,30 @@ def _add_inference_args(parser):
help="Host<->device bandwidth for the KV offload tier in GB/s. "
"PCIe 5 x16 is ~64; a cache-coherent host link is ~900. Default: 64.",
)
parser.add_argument(
"--workload-resident-tokens",
type=int,
default=None,
help="Mean KV tokens one request of the target workload holds while "
"resident, which sets how many requests the pool can admit at once and "
"so where TTFT stops being a service time. For a spread of lengths pass "
"E[L^2]/E[L], not the mean: a request holds the pool for a time "
"proportional to its length, so the resident set is length-biased. Do "
"not discount it by the prefix-cache hit rate -- a cache hit skips "
"prefill compute, it does not free the blocks. Default: the configured "
"context (input + output/2), exact for a fixed-length workpoint.",
)
parser.add_argument(
"--kv-pool-tokens",
type=int,
default=None,
help="Size of the KV pool the engine allocated, in tokens, as it "
'reports at startup (vLLM\'s "GPU KV cache size", SGLang\'s "KV Cache '
'is allocated. #tokens"). The admission bound is evaluated against '
"this instead of against the memory model's estimate, which is what "
"you want when validating against a deployment that is already "
"running. Default: unset, and the pool is predicted.",
)
# ---- Feature B: custom collective ops ----
coll = parser.add_argument_group("inference collectives (feature B)")
coll.add_argument(
Expand Down Expand Up @@ -616,6 +640,17 @@ def _add_inference_args(parser):
"parallelism replicates their compressed KV latent instead of sharding "
"it, so only splitting by request shrinks the cache a rank holds.",
)
par.add_argument(
"--decode-context-parallel-size",
type=int,
default=None,
help="Split each sequence's context across this many ranks during "
"decode (vLLM --decode-context-parallel-size, SGLang --dcp-size). The "
"other way to shrink an MLA cache: a rank stores context/N of every "
"request rather than all of it, so the replica holds N times as many "
"tokens. Unlike --attention-dp-size it shrinks a single request's "
"footprint, so it raises the concurrency ceiling even at batch 1.",
)
# ---- Feature A: prefill/decode disaggregation ----
dis = parser.add_argument_group("inference disaggregation (feature A)")
dis.add_argument(
Expand Down Expand Up @@ -790,6 +825,20 @@ def _add_inference_args(parser):
"only (not throughput). A property of the serving stack, so it has no default "
"worth guessing. Default 0.",
)
serv.add_argument(
"--uncached-prompt-latency-us",
type=float,
default=None,
help="TTFT latency per prompt token the prefix cache did not serve (us/token). "
"Added to TTFT and end-to-end latency only (not throughput). Default 0.",
)
serv.add_argument(
"--uncached-prompt-latency-max-tokens",
type=int,
default=None,
help="Uncached tokens past which --uncached-prompt-latency-us stops growing. "
"Default 0 (no cap).",
)
serv.add_argument(
"--prefill-rate-us-per-token",
type=float,
Expand Down Expand Up @@ -853,6 +902,14 @@ def _add_inference_args(parser):
"prefill-chunk + concurrent-decode tokens per step; oversized steps split, "
"raising TPOT. Default: 0 (unlimited).",
)
serv.add_argument(
"--max-num-seqs",
type=int,
default=None,
help="Engine cap on running sequences (vLLM --max-num-seqs, SGLang "
"--max-running-requests). Clients beyond it queue at the server. "
"Default: 0 (bounded by client concurrency only).",
)
# ---- Offered load / request rate (open-loop arrivals) ----
serv.add_argument(
"--request-rate",
Expand Down Expand Up @@ -939,6 +996,75 @@ def _add_inference_args(parser):
"fixed-concurrency run is only a few waves long and the opening burst "
"carries the highest TTFT of any request in it. Default: 0.1.",
)
serv.add_argument(
"--des-client-think-ms",
type=float,
default=0.0,
help="DES: how long a closed-loop client waits after its last token "
"before issuing its next request. Zero re-issues immediately, which "
"is a load generator saturating the server and not what an agentic "
"harness does: a turn ends, the agent runs a tool, and the next turn "
"is sent when that returns. It is a property of the workload, like "
"the prompt lengths and the prefix reuse, and leaving it out does not "
"leave the replay neutral -- it holds every client permanently "
"resident. On the AgentX ladders the hardware's clients are idle for "
"50-74% of each turn's cycle below saturation and ~0% above it, so "
"omitting it overstates throughput several-fold at the low rungs and "
"not at all at the high ones, which bends the curve rather than "
"shifting it. A Mooncake trace's per-request ``think_ms`` overrides "
"it request by request.",
)
serv.add_argument(
"--des-client-idle-cap-ms",
type=float,
default=0.0,
help="DES: longest the engine may sit with no request in flight before "
"pending client timers are pulled forward (0 = no cap). Mirrors a "
"replay harness's system-idle guard (AIPerf AgentX: 10 s), which "
"shortens a single client's idle time far below the trace's gaps.",
)
serv.add_argument(
"--des-admit-backlog-only",
action="store_true",
help="DES: measure the longest-prefix admission window against the "
"backlog rather than the client count, so that below what the KV pool "
"holds -- where nothing is queued and there is no choice for a policy "
"to make -- admission is the order the lanes arrived in. This is the "
"correct account of what a scheduler can reorder, and it is the band "
"where measured reuse collapses (kimik3 at C=14 thrashed to 0.67 "
"against 0.93 modelled), but on the AgentX corpus it improves ITL "
"ordering by one model-engine pair and costs one on throughput, so it "
"is off by default.",
)
serv.add_argument(
"--des-warmup-requests",
type=int,
default=0,
help="DES: exclude this many requests, in the order the clients issued "
"them, from the reported latency distribution. Prefer this to "
"--des-warmup-frac for a closed loop: all C clients fire at once, so "
"the first C requests queue against each other and wait far longer "
"than anything after them, and dropping the earliest *completions* "
"keeps every one of them -- a request that waited 40s for a slot is "
"among the last to finish. The harness excludes the same transient by "
"advancing each lane AIPERF_WARMUP_REQUESTS_PER_LANE requests and "
"draining before it starts profiling, so the matching value is that "
"many per client. Default: 0 (use --des-warmup-frac).",
)
serv.add_argument(
"--des-duration-s",
type=float,
default=0.0,
help="DES: stop the run after this many seconds of simulated time and "
"report over whatever completed, instead of running to "
"--des-num-requests. This is how a fixed-concurrency harness is "
"actually bounded, and under a closed loop the two are not "
"interchangeable: a request budget divided among C lanes fixes how "
"far each lane walks into its conversation, and an agentic turn's "
"prompt grows with its position, so a budget-bounded run offers "
"different prompt lengths at each concurrency than the measured run "
"it is compared against. Default: 0 (run to the request count).",
)
serv.add_argument(
"--des-seed",
type=int,
Expand Down Expand Up @@ -1064,6 +1190,34 @@ def _add_inference_args(parser):
help="DES: per-instance KV block-cache capacity (blocks kept resident, "
"LRU-evicted under pressure). 0 = unbounded within the run. Default: 0.",
)
serv.add_argument(
"--des-whole-context-residency",
action="store_true",
help="DES: charge a running request the whole context it attends to, "
"not just the part of it that was not already cached. The discount is "
"an account of one physical copy, and it is only that where the "
"sharers are running at the same time. Under a closed loop on "
"conversational traffic they are not: each client has one turn in "
"flight, reuse is a conversation hitting its own previous turn, and "
"the turns running together belong to different conversations whose "
"histories do not overlap. Discounted anyway, a pool holding five of "
"these conversations reports room for hundreds, never fills, and so "
"never shows reuse falling away or the queue that follows it.",
)
serv.add_argument(
"--des-cache-shares-pool",
action="store_true",
help="DES: size the prefix cache to what the running requests leave "
"free, instead of to the whole pool. A request is already not charged "
"for the blocks it hit on, because something else is holding them; "
"that something is this cache, and giving it the pool as well books "
"the same memory twice. Where the pool is roomy the correction is "
"nothing, and where it is not it is the whole behaviour: a pool with "
"space for five of these conversations cannot also hold their history, "
"so reuse falls away as concurrency climbs and TTFT leaves the scale "
"the low rungs were on. Double-booked, the replay reports reuse near "
"its ceiling and a flat TTFT straight through that.",
)
serv.add_argument(
"--des-mooncake-trace",
type=str,
Expand All @@ -1084,6 +1238,17 @@ def _add_inference_args(parser):
help="Attention kernel library (ROCm). Representative compute multiplier "
"vs the Triton baseline. Default: engine default (1.0).",
)
kern.add_argument(
"--serving-engine",
type=str,
default=None,
help="Engine this projection is about (vllm, sglang, atom, or a build "
"such as mori-sglang). Simulate mode is unaffected -- it is analytical "
"and returns the same number whichever engine is named. This is a "
"regime axis, so it decides which measured anchor benchmark mode may "
"calibrate from: an anchor harvested on another engine is refused "
"rather than substituted. Default: unset (matches as before).",
)
kern.add_argument(
"--sparse-attention-topk",
type=int,
Expand All @@ -1092,6 +1257,18 @@ def _add_inference_args(parser):
"query. Attention scales toward topk/context for long contexts. "
"Default: 0 (dense).",
)
kern.add_argument(
"--sparse-indexer-cost-scale",
type=float,
default=None,
help="What the sparse indexer's top-k selection costs on this stack "
"relative to a fused kernel. The arithmetic is the model's, the kernel "
"is the serving stack's: the same selection runs fused through aiter "
"on gfx950 and unfused through Torch on gfx942. 1.0 prices the fused "
"path; raise it to charge a stack serving without one. Only reaches "
"models priced from an indexer geometry (GLM-5.2, MiniMax-M3), not "
"ones carrying a per-layer compression schedule. Default: 1.0.",
)
kern.add_argument(
"--sliding-window",
type=int,
Expand Down
21 changes: 21 additions & 0 deletions infera/projection/configs/models/megatron/glm5.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,27 @@ extends:
# declared below so tooling does not have to infer it from this comment.
sparse_attention_topk: 2048

# IndexShare is a memory property as much as a compute one. The indexer stores
# a 128-wide key per token (``index_head_dim: 128``), but only on the layers
# that compute one: ``indexer_types`` marks 21 of the 78 layers ``full`` and
# the other 57 ``shared``, and a shared layer reuses the keys of the full
# layer above it rather than storing its own. GLM keeps them fp8, unlike
# MiniMax-M3's bf16, so here the index cache is a 6% correction rather than a
# third of the footprint: against vLLM at TP8 this predicts 47,616 bytes per
# token with an fp8 KV cache where the engine allocates 47,805, and 92,544
# with a bf16 cache where it allocates 92,815.
sparse_index_head_dim: 128
sparse_index_layers: 21
sparse_index_dtype: fp8

# The indexer's query heads (HF ``index_n_heads: 32``). Storage is unaffected
# -- the 128-wide key above is shared across them -- but the arithmetic is
# not, because each head scores the whole context. Without it GLM has no
# indexer term at all and prices back to a floored top-k, which on this
# corpus is a constant 0.15 of dense from 136k to 365k: an over-read of the
# selection it names and a free pass on the indexing it does not.
sparse_index_n_heads: 32

tokenizer_type: HuggingFaceTokenizer
##### TODO update to GLM-5 tokenizer
tokenizer_model: THUDM/glm-4-9b # GLM-5 tokenizer requires newer transformers; use GLM-4 for mock_data testing
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
bases:
extends:
- mamba_base.yaml

# Mamba 370M configuration
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
bases:
extends:
- language_model.yaml

# Mamba-specific configuration
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
bases:
extends:
- language_model.yaml

# MiniMax-M2.5 Model (GQA + MoE)
Expand Down
21 changes: 20 additions & 1 deletion infera/projection/configs/models/megatron/minimax_m3.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
bases:
extends:
- language_model.yaml

# MiniMax-M3 (GQA + sparse attention + MoE)
Expand All @@ -16,6 +16,25 @@ bases:
# as a serving default rather than as an architectural dimension.
sparse_attention_topk: 2048

# Choosing those 16 blocks has a memory cost, and on this model it is a third
# of the footprint. The indexer stores a 128-wide key per token
# (``sparse_index_dim: 128``) on each of the 57 layers where
# ``sparse_attention_freq`` is set -- the first three are dense and have no
# indexer -- and keeps them in bf16, so unlike the KV cache they do not shrink
# when it is quantized. Against vLLM at TP8 this predicts 29,952 bytes per
# token with an fp8 cache where the engine allocates 30,023, and 45,312 with a
# bf16 cache where it allocates 45,398.
sparse_index_head_dim: 128
sparse_index_layers: 57
sparse_index_dtype: bf16

# The indexer's query heads (HF ``sparse_num_index_heads: 4``). Eight times
# narrower than GLM-5.2's 32 against attention layers that are themselves
# four times cheaper -- GQA over 64 heads of 128 rather than absorbed MLA --
# so M3's indexer lands at a comparable fraction of its own attention from
# very different parts. Both are invisible without this field.
sparse_index_n_heads: 4

tokenizer_type: HuggingFaceTokenizer
tokenizer_model: MiniMaxAI/MiniMax-M3

Expand Down
Loading
Loading