Skip to content

MInf: hand compact prompt ids and media tensors to the payload stager - #20

Closed
lauradang wants to merge 119 commits into
tdene:mainfrom
lauradang:laurad/pr7015-multimodal-capture
Closed

lauradang wants to merge 119 commits into
tdene:mainfrom
lauradang:laurad/pr7015-multimodal-capture

Conversation

@lauradang

Copy link
Copy Markdown
Collaborator

What does this PR do?

Stacked on NVIDIA#7015 (tde/ledger_capture), which is this PR's base branch.

Token capture of multimodal requests needs two things the offloaded payload did not carry:

  • compact_prompt_token_ids: the pre-expansion prompt the chat endpoint tokenized. A RequestPromptPreparer runs before add_request, so it can only splice a previous turn in compact space; splicing the expanded previous turn makes _build_vlm_request expand it a second time and reject the request on the placeholder count. The preparer docstring now states the compact-space contract, and a completions test pins that failure.
  • media_tensors: host copies of the vision encoder's inputs (imgs as packed patches, imgs_sizes, num_frames / num_tiles), so the consumer can train on exactly the pixels the policy generated against. They live on DynamicInferenceRequest.media_tensors, survive checkpoint() and merge(), and are dropped from the wire.

Both fields are None for text-only requests. Consumers: lauradang/Gym and lauradang/RL PRs stacked on NVIDIA-NeMo/Gym#2823 and NVIDIA-NeMo/RL#4129.

Pre-checks

  • Unit tests added (test_inference_request.py, test_multimodal_entrypoints.py, request lifecycle policy table)
  • Typing on new fields
  • black and isort run on the changed files

🤖 Generated with Claude Code

fwerkor and others added 30 commits September 3, 2026 11:49
Signed-off-by: Cao Yuhang <caoyuhang@fwerkor.com>
Signed-off-by: Castronaut <3116935098@qq.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
…819] (NVIDIA#7024)

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Yaniv Galron <ygalron@nvidia.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
…est (NVIDIA#7041)

Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: Chase Block <cblock@nvidia.com>
Signed-off-by: Sudhakar Singh <sudhakars@nvidia.com>
Signed-off-by: Ravi Ghadia <rghadia@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com>
Co-authored-by: Sudhakar Singh <sudhakars@nvidia.com>
Co-authored-by: Ravi Ghadia <rghadia@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: William Dykas <wdykas@nvidia.com>
Signed-off-by: shanmugamr1992 <shanmugamr1992@gmail.com>
Co-authored-by: shanmugamr1992 <shanmugamr1992@gmail.com>
…erminism (NVIDIA#7069)

Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
…sformerConfig (NVIDIA#5756)

Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Co-authored-by: Charlie Truong <chtruong@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
…biencoder_args (NVIDIA#6975)

Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: Tom Long <tolong@nvidia.com>
Signed-off-by: ilml <tolong@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Haoran Zhang <haoranz@nvidia.com>
Signed-off-by: Jack Chang <jianbinc@nvidia.com>
Signed-off-by: jeffnvidia <jmahou@nvidia.com>
Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com>
Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
faradawn and others added 20 commits September 14, 2026 17:28
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
…IA#6956)

Signed-off-by: Prajwal Singhania <psinghania@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
…_chunked_prefill (NVIDIA#7334)

Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Jimmy Zhang <jiemingz@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Resolve conflicts with main:
- inference_request.py: combine main's prompt-view drop (prompt, compact,
  remaining) with payload-offload field stripping and stage metadata.
- dynamic_engine.py: apply payload staging on main's flat-request path.
- test_inference_request.py: merge the parametrized return_prompt_tokens cases.
- multimodal/utils.py: removed on main by NVIDIA#7326; drop the einops guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
The per-request dict that rides with a SUBMIT_REQUEST to the engine's
payload stager and prompt preparer was called request_metadata. That name
already denotes the per-request sampling tensors on DynamicInferenceContext
(request_metadata, active_request_metadata, request_metadata_types), and it
sat in a wire list whose other fields are request metadata as well. Rename
it end to end: REST body key, InferenceClient kwargs, coordinator handler,
engine, and the DynamicInferenceRequest field. The sampling-tensor names are
unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Ship the chat-template render of the conversation through its last
assistant message (template_prefix_token_ids) and the EOS id in
offload_params instead of a precomputed splice boundary. The engine-side
RequestPromptPreparer (NeMo RL) applies its own replace_prefix_tokens, so
Megatron and NeMo RL no longer run two different boundary algorithms on one
request. _replace_prefix_tokens is restored to the main version and keeps
serving clients that echo prior-turn token ids and install no preparer.

Remove required_prefix_token_ids: the preparer resolves the exact prefix
itself, so the endpoint no longer needs to accept one.

Decide the prefix source once, up front: token ids on the last assistant
message (prefix replaced here) or offload_params (prefix replaced in the
engine). Supplying both is rejected with 400.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Follow the engine API after the merge with main: send flat requests to the coordinator, read step_modern()['finished_requests'], accept the offload_params keyword in add_request stubs, and unpack the 5-field SUBMIT_REQUEST metadata frame. Classify the new offload_params, payload_offloaded and payload_stage_metadata request fields in the lifecycle policy table. Add class-level None defaults for payload_stager and prompt_preparer so engines built via __new__ in tests are readable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Use the module logger instead of the root logger when a prompt preparer raises.
- Pack the prepared prompt and params inside the try so an unserializable preparer result falls back to the original frame stamped with the preparation error, failing the request as PromptPreparationError instead of exiting MP rank 0.
- Extract _apply_prompt_preparation_error and call it from both add_request branches.
- Release VLM media registered by _build_vlm_request in _handle_failed_request so a request that fails before admission does not leak it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
C12: move offload_params out of the metadata frame into a fifth client frame that the coordinator forwards verbatim (never decoded), so the frame it unpacks and repacks per request stays bounded; the engine reads it from its own frame at admission and the preparer rewrites only the prompt and offload frames.
C11: schedule_requests logs a warning and drops a SUBMIT_REQUEST with the wrong frame or metadata-field count instead of raising, so one version-skewed client cannot take every MP rank down; the skip is collective because all ranks see the same broadcast list.
C13: _prepare_submit_request_message returns the message untouched, without decoding any frame, when the offload frame is msgpack nil, since the preparer only has work when the client sent params.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…contract

- serialize(): an offloaded reply never carries the prompt tensors (the stager holds the prompt ids); otherwise they stay opt-in via return_prompt_tokens and are dropped when sampling_params is None, as on main.
- _send_requests_to_coordinator(): build FinishedRequestRecord once per completed request and share it between the ledger and _serialize_finished_request().
- OffloadedRequestPayload.from_request(): document the prompt_tokens host-copy side effect.
- payload_offloaded / payload_stage_metadata: reword the field comment; they are wire-only keys kept for deserialize().
- checkpoint()/merge(): share offload_params instead of deep-copying; nothing mutates it in place.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
The finished-request rename left two references to the old merged record in the GPU-only offload test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… in offload_params

E19: /v1/completions now collects each reply's payload_stage_metadata, rejects conflicting values across a batch, and merges it into the top-level response body with the same reserved-field check as /v1/chat/completions; the shared logic lives in endpoints/common.py (collect_stage_metadata, attach_stage_metadata). The endpoint also accepts and forwards offload_params, which it previously ignored.
E20: both endpoints reject client-supplied offload_params whose top-level keys start with '_' (HTTP 400) via endpoints.common.validate_offload_params, so the engine-owned _request_prompt_preparation_error field cannot be forged from a request body.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…load stager

OffloadedRequestPayload gains compact_prompt_token_ids (the pre-expansion
prompt the endpoint tokenized) and media_tensors (host copies of the vision
encoder's inputs: packed-patch imgs, imgs_sizes, num_frames / num_tiles). Both
are None for text-only requests. The tensors live on
DynamicInferenceRequest.media_tensors, survive checkpoint() and merge(), and
are dropped from the wire.

Without the compact form an external RequestPromptPreparer can only splice the
expanded previous turn into the compact new render, which _build_vlm_request
then expands a second time and rejects on the placeholder count. The preparer
docstring now states the compact-space contract, and a completions test pins
the admission failure an already-expanded prefix produces.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@lauradang

Copy link
Copy Markdown
Collaborator Author

/claude strict-review

@lauradang lauradang self-assigned this Sep 18, 2026
lauradang added a commit to lauradang/RL that referenced this pull request Sep 18, 2026
- ChainPrefixCache now requires TQTokenSource.fetch_prefix_chains. The
  duck-typed fallback served only a test double and, if ever taken, would
  have silently collapsed the compact chain onto the expanded one (the
  double-expansion this PR fixes). Test doubles implement the method.
- Extract the media column codec into media_field_dict /
  row_to_media_tensors, matching the module's _row_to_* decoders, so
  stage_media / fetch_media are one-line delegations and the codec
  round-trips without a TransferQueue.
- Fix comments still describing the terminal-row design: media is staged
  as per-call deltas and concatenated along the chain. Correct the
  _media_mismatch docstring (the summary carries no embedding count).
  Reconcile the design doc's "outside the token rows" / "ride the call
  row" wording and name the unpinned tdene/Megatron-LM#20 and Gym
  staging/media.py dependencies. Sort the tq_token_sink import block.
- Tests: slice_media_tensors integrity guards (no geometry, non-dividing
  patch count, mid-patch boundary), media_columns_missing rejection after
  a swallowed stage_media failure, codec round-trip (f32 / bf16 / video /
  tiled), Gym extras-key and adapter-attribute parity, text runs register
  no media columns, expanded-space mask/logprob alignment, and an @mcore
  pin guard asserting OffloadedRequestPayload exposes media_tensors and
  compact_prompt_token_ids (red until the Megatron-Bridge pointer includes
  tdene/Megatron-LM#20).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
@lauradang
lauradang changed the base branch from tde/ledger_capture to main September 21, 2026 16:50
@lauradang

Copy link
Copy Markdown
Collaborator Author

Superseded by NVIDIA#7557, which carries the same commit rebased onto NVIDIA main now that NVIDIA#7015 has merged.

@lauradang lauradang closed this Sep 21, 2026
lauradang added a commit to lauradang/RL that referenced this pull request Sep 29, 2026
- ChainPrefixCache now requires TQTokenSource.fetch_prefix_chains. The
  duck-typed fallback served only a test double and, if ever taken, would
  have silently collapsed the compact chain onto the expanded one (the
  double-expansion this PR fixes). Test doubles implement the method.
- Extract the media column codec into media_field_dict /
  row_to_media_tensors, matching the module's _row_to_* decoders, so
  stage_media / fetch_media are one-line delegations and the codec
  round-trips without a TransferQueue.
- Fix comments still describing the terminal-row design: media is staged
  as per-call deltas and concatenated along the chain. Correct the
  _media_mismatch docstring (the summary carries no embedding count).
  Reconcile the design doc's "outside the token rows" / "ride the call
  row" wording and name the unpinned tdene/Megatron-LM#20 and Gym
  staging/media.py dependencies. Sort the tq_token_sink import block.
- Tests: slice_media_tensors integrity guards (no geometry, non-dividing
  patch count, mid-patch boundary), media_columns_missing rejection after
  a swallowed stage_media failure, codec round-trip (f32 / bf16 / video /
  tiled), Gym extras-key and adapter-attribute parity, text runs register
  no media columns, expanded-space mask/logprob alignment, and an @mcore
  pin guard asserting OffloadedRequestPayload exposes media_tensors and
  compact_prompt_token_ids (red until the Megatron-Bridge pointer includes
  tdene/Megatron-LM#20).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.