Conversation
Signed-off-by: Cao Yuhang <caoyuhang@fwerkor.com> Signed-off-by: Castronaut <3116935098@qq.com> Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
…819] (NVIDIA#7024) Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Teodor-Dumitru Ene <teodord.ene@gmail.com>
Signed-off-by: Tom Long <tolong@nvidia.com>
Signed-off-by: Yaniv Galron <ygalron@nvidia.com>
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
…est (NVIDIA#7041) Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
Signed-off-by: Lawrence McAfee <lmcafee@nvidia.com>
Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: Chase Block <cblock@nvidia.com> Signed-off-by: Sudhakar Singh <sudhakars@nvidia.com> Signed-off-by: Ravi Ghadia <rghadia@nvidia.com> Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com> Co-authored-by: Siddhartha Raman Sundara Raman <sraman@nvidia.com> Co-authored-by: Sudhakar Singh <sudhakars@nvidia.com> Co-authored-by: Ravi Ghadia <rghadia@nvidia.com> Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Co-authored-by: Fei Wu <33940270+YangFei1990@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: William Dykas <wdykas@nvidia.com> Signed-off-by: shanmugamr1992 <shanmugamr1992@gmail.com> Co-authored-by: shanmugamr1992 <shanmugamr1992@gmail.com>
…erminism (NVIDIA#7069) Signed-off-by: Jingyue Wu <jingyuew@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
…sformerConfig (NVIDIA#5756) Signed-off-by: Rui Zhu <rui.zhu.rz399@yale.edu> Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com> Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Co-authored-by: Charlie Truong <chtruong@nvidia.com> Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
…biencoder_args (NVIDIA#6975) Signed-off-by: Anai-Guo <antai12232931@outlook.com> Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: Tom Long <tolong@nvidia.com> Signed-off-by: ilml <tolong@nvidia.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Zhiyu Li <zhiyul@nvidia.com>
Signed-off-by: Haoran Zhang <haoranz@nvidia.com>
Signed-off-by: Jack Chang <jianbinc@nvidia.com>
Signed-off-by: jeffnvidia <jmahou@nvidia.com> Signed-off-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com> Co-authored-by: svcnvidia-nemo-ci <svcnvidia-nemo-ci@nvidia.com>
Signed-off-by: dimapihtar <dpykhtar@nvidia.com> Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
Signed-off-by: Deyu Fu <deyuf@nvidia.com> Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com> Co-authored-by: Dmytro Pykhtar <37850217+dimapihtar@users.noreply.github.com>
…IA#6956) Signed-off-by: Prajwal Singhania <psinghania@nvidia.com>
Signed-off-by: wdykas <wdykas@nvidia.com>
…_chunked_prefill (NVIDIA#7334) Signed-off-by: Charlie Truong <chtruong@nvidia.com>
Signed-off-by: Philip Petrakian <ppetrakian@nvidia.com>
…7293) Signed-off-by: Shiqing Fan <shiqingf@nvidia.com>
Signed-off-by: Jimmy Zhang <jiemingz@nvidia.com>
Signed-off-by: janEbert <janpabloe@nvidia.com>
Signed-off-by: tongliu <tongliu@nvidia.com>
Resolve conflicts with main: - inference_request.py: combine main's prompt-view drop (prompt, compact, remaining) with payload-offload field stripping and stage metadata. - dynamic_engine.py: apply payload staging on main's flat-request path. - test_inference_request.py: merge the parametrized return_prompt_tokens cases. - multimodal/utils.py: removed on main by NVIDIA#7326; drop the einops guard. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
The per-request dict that rides with a SUBMIT_REQUEST to the engine's payload stager and prompt preparer was called request_metadata. That name already denotes the per-request sampling tensors on DynamicInferenceContext (request_metadata, active_request_metadata, request_metadata_types), and it sat in a wire list whose other fields are request metadata as well. Rename it end to end: REST body key, InferenceClient kwargs, coordinator handler, engine, and the DynamicInferenceRequest field. The sampling-tensor names are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Ship the chat-template render of the conversation through its last assistant message (template_prefix_token_ids) and the EOS id in offload_params instead of a precomputed splice boundary. The engine-side RequestPromptPreparer (NeMo RL) applies its own replace_prefix_tokens, so Megatron and NeMo RL no longer run two different boundary algorithms on one request. _replace_prefix_tokens is restored to the main version and keeps serving clients that echo prior-turn token ids and install no preparer. Remove required_prefix_token_ids: the preparer resolves the exact prefix itself, so the endpoint no longer needs to accept one. Decide the prefix source once, up front: token ids on the last assistant message (prefix replaced here) or offload_params (prefix replaced in the engine). Supplying both is rejected with 400. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Follow the engine API after the merge with main: send flat requests to the coordinator, read step_modern()['finished_requests'], accept the offload_params keyword in add_request stubs, and unpack the 5-field SUBMIT_REQUEST metadata frame. Classify the new offload_params, payload_offloaded and payload_stage_metadata request fields in the lifecycle policy table. Add class-level None defaults for payload_stager and prompt_preparer so engines built via __new__ in tests are readable. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
- Use the module logger instead of the root logger when a prompt preparer raises. - Pack the prepared prompt and params inside the try so an unserializable preparer result falls back to the original frame stamped with the preparation error, failing the request as PromptPreparationError instead of exiting MP rank 0. - Extract _apply_prompt_preparation_error and call it from both add_request branches. - Release VLM media registered by _build_vlm_request in _handle_failed_request so a request that fails before admission does not leak it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
C12: move offload_params out of the metadata frame into a fifth client frame that the coordinator forwards verbatim (never decoded), so the frame it unpacks and repacks per request stays bounded; the engine reads it from its own frame at admission and the preparer rewrites only the prompt and offload frames. C11: schedule_requests logs a warning and drops a SUBMIT_REQUEST with the wrong frame or metadata-field count instead of raising, so one version-skewed client cannot take every MP rank down; the skip is collective because all ranks see the same broadcast list. C13: _prepare_submit_request_message returns the message untouched, without decoding any frame, when the offload frame is msgpack nil, since the preparer only has work when the client sent params. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
…contract - serialize(): an offloaded reply never carries the prompt tensors (the stager holds the prompt ids); otherwise they stay opt-in via return_prompt_tokens and are dropped when sampling_params is None, as on main. - _send_requests_to_coordinator(): build FinishedRequestRecord once per completed request and share it between the ledger and _serialize_finished_request(). - OffloadedRequestPayload.from_request(): document the prompt_tokens host-copy side effect. - payload_offloaded / payload_stage_metadata: reword the field comment; they are wire-only keys kept for deserialize(). - checkpoint()/merge(): share offload_params instead of deep-copying; nothing mutates it in place. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
The finished-request rename left two references to the old merged record in the GPU-only offload test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
… in offload_params E19: /v1/completions now collects each reply's payload_stage_metadata, rejects conflicting values across a batch, and merges it into the top-level response body with the same reserved-field check as /v1/chat/completions; the shared logic lives in endpoints/common.py (collect_stage_metadata, attach_stage_metadata). The endpoint also accepts and forwards offload_params, which it previously ignored. E20: both endpoints reject client-supplied offload_params whose top-level keys start with '_' (HTTP 400) via endpoints.common.validate_offload_params, so the engine-owned _request_prompt_preparation_error field cannot be forged from a request body. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…load stager OffloadedRequestPayload gains compact_prompt_token_ids (the pre-expansion prompt the endpoint tokenized) and media_tensors (host copies of the vision encoder's inputs: packed-patch imgs, imgs_sizes, num_frames / num_tiles). Both are None for text-only requests. The tensors live on DynamicInferenceRequest.media_tensors, survive checkpoint() and merge(), and are dropped from the wire. Without the compact form an external RequestPromptPreparer can only splice the expanded previous turn into the compact new render, which _build_vlm_request then expands a second time and rejects on the placeholder count. The preparer docstring now states the compact-space contract, and a completions test pins the admission failure an already-expanded prefix produces. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Closed
5 of 6 tasks
Collaborator
Author
|
/claude strict-review |
lauradang
added a commit
to lauradang/RL
that referenced
this pull request
Sep 18, 2026
- ChainPrefixCache now requires TQTokenSource.fetch_prefix_chains. The duck-typed fallback served only a test double and, if ever taken, would have silently collapsed the compact chain onto the expanded one (the double-expansion this PR fixes). Test doubles implement the method. - Extract the media column codec into media_field_dict / row_to_media_tensors, matching the module's _row_to_* decoders, so stage_media / fetch_media are one-line delegations and the codec round-trips without a TransferQueue. - Fix comments still describing the terminal-row design: media is staged as per-call deltas and concatenated along the chain. Correct the _media_mismatch docstring (the summary carries no embedding count). Reconcile the design doc's "outside the token rows" / "ride the call row" wording and name the unpinned tdene/Megatron-LM#20 and Gym staging/media.py dependencies. Sort the tq_token_sink import block. - Tests: slice_media_tensors integrity guards (no geometry, non-dividing patch count, mid-patch boundary), media_columns_missing rejection after a swallowed stage_media failure, codec round-trip (f32 / bf16 / video / tiled), Gym extras-key and adapter-attribute parity, text runs register no media columns, expanded-space mask/logprob alignment, and an @mcore pin guard asserting OffloadedRequestPayload exposes media_tensors and compact_prompt_token_ids (red until the Megatron-Bridge pointer includes tdene/Megatron-LM#20). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Merged
5 of 6 tasks
Collaborator
Author
|
Superseded by NVIDIA#7557, which carries the same commit rebased onto NVIDIA main now that NVIDIA#7015 has merged. |
lauradang
added a commit
to lauradang/RL
that referenced
this pull request
Sep 29, 2026
- ChainPrefixCache now requires TQTokenSource.fetch_prefix_chains. The duck-typed fallback served only a test double and, if ever taken, would have silently collapsed the compact chain onto the expanded one (the double-expansion this PR fixes). Test doubles implement the method. - Extract the media column codec into media_field_dict / row_to_media_tensors, matching the module's _row_to_* decoders, so stage_media / fetch_media are one-line delegations and the codec round-trips without a TransferQueue. - Fix comments still describing the terminal-row design: media is staged as per-call deltas and concatenated along the chain. Correct the _media_mismatch docstring (the summary carries no embedding count). Reconcile the design doc's "outside the token rows" / "ride the call row" wording and name the unpinned tdene/Megatron-LM#20 and Gym staging/media.py dependencies. Sort the tq_token_sink import block. - Tests: slice_media_tensors integrity guards (no geometry, non-dividing patch count, mid-patch boundary), media_columns_missing rejection after a swallowed stage_media failure, codec round-trip (f32 / bf16 / video / tiled), Gym extras-key and adapter-attribute parity, text runs register no media columns, expanded-space mask/logprob alignment, and an @mcore pin guard asserting OffloadedRequestPayload exposes media_tensors and compact_prompt_token_ids (red until the Megatron-Bridge pointer includes tdene/Megatron-LM#20). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Stacked on NVIDIA#7015 (
tde/ledger_capture), which is this PR's base branch.Token capture of multimodal requests needs two things the offloaded payload did not carry:
compact_prompt_token_ids: the pre-expansion prompt the chat endpoint tokenized. ARequestPromptPreparerruns beforeadd_request, so it can only splice a previous turn in compact space; splicing the expanded previous turn makes_build_vlm_requestexpand it a second time and reject the request on the placeholder count. The preparer docstring now states the compact-space contract, and a completions test pins that failure.media_tensors: host copies of the vision encoder's inputs (imgsas packed patches,imgs_sizes,num_frames/num_tiles), so the consumer can train on exactly the pixels the policy generated against. They live onDynamicInferenceRequest.media_tensors, survivecheckpoint()andmerge(), and are dropped from the wire.Both fields are
Nonefor text-only requests. Consumers: lauradang/Gym and lauradang/RL PRs stacked on NVIDIA-NeMo/Gym#2823 and NVIDIA-NeMo/RL#4129.Pre-checks
test_inference_request.py,test_multimodal_entrypoints.py, request lifecycle policy table)🤖 Generated with Claude Code