Skip to content

feat(token-capture): carry media through Megatron capture rollouts - #4192

Closed
lauradang wants to merge 25 commits into
NVIDIA-NeMo:mainfrom
lauradang:laurad/pr4129-multimodal-capture
Closed

lauradang wants to merge 25 commits into
NVIDIA-NeMo:mainfrom
lauradang:laurad/pr4129-multimodal-capture

Conversation

@lauradang

Copy link
Copy Markdown
Contributor

What does this PR do ?

Stacked on #4129 (laurad/minf-rollout-checkpointing); the diff shows that PR's commits until it merges. Only the top commit is new here.

Token capture on the Megatron backend was text-only. Receipt rows carried no pixel_values, so a VLM capture run trained image-blind, and every multi-turn VLM rollout failed at admission because the preparer spliced the expanded previous turn into the compact new render and Megatron expanded it twice.

  • Compact chains. STAGING_FIELDS gains compact_token_ids_delta / compact_len. TQTokenSink pops Gym's compact delta out of extras into that column (like routed_experts); TQTokenSource.fetch_prefix_chains returns both token spaces, with compact falling back to expanded for text calls.
  • Preparer. TQMegatronPromptPreparer splices with the compact chain, hands Gym the expanded chain to verify the engine prompt against, and records offload_params["ng_capture_minf"]["compact_prev_len"] for the stager.
  • Media on the call row. TQMegatronTokenStager derives the digest-covered media geometry from MInf's media_tensors for Gym's adapter, then stages the tensors themselves as extra columns on the same staging key (MEDIA_STAGING_FIELDS), the way routed experts ride the row.
  • Finalizer. RolloutReassembler reads the terminal call's media columns, checks them against the staged geometry (media_columns_missing / media_mismatch), and publishes the engine's packed patches as pixel_values with imgs_sizes and num_frames. The Megatron-Bridge Omni model accepts the [total_patches, C*P*P] layout directly, so training projects exactly the pixels the policy generated against.
  • Setup. Registers the media columns for multimodal runs; rejects capture with a multimodal policy on the vLLM backend (that path stages the pre-processor prompt and carries no media) and with grpo.deduplicate_multimodal_data.

docs/design-docs/token-capture-ledger.md gains a "Multimodal rollouts" section.

Dependencies

Testing

  • tests/unit/data_plane/test_tq_token_sink.py and test_rollout_reassembler.py against the in-process TransferQueue simple backend: 56 passed
  • tests/unit/single_controller/test_setup.py -k "multimodal or megatron_token_capture": 7 passed
  • Not yet run: the end-to-end Nemotron Omni + Gym + capture functional test on GPU, which needs the Megatron-Bridge and Gym re-pins.

Before your PR is "Ready for review"

  • Did you write any new necessary tests?
  • Did you run the unit tests locally?
  • Did you add or update any necessary documentation?

🤖 Generated with Claude Code

lauradang and others added 25 commits September 14, 2026 13:09
Signed-off-by: Laura Dang <laurad@nvidia.com>
…very.sh

Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… cache

Extract the vLLM worker's _fetch_chain_prefix and _resolve_admission_prefix
into tq_token_sink.py as ChainPrefixCache and resolve_admission_prefix, and
make the worker methods one-line delegates. TQMegatronPromptPreparer now
resolves staging chains through the same pair, so both backends share one
cached TQ read (256 entries keyed by the chain's last staging key, deepest
cached key bounds the fetch to the uncached suffix).

The preparer reads the splice boundary from the request-metadata keys the
Megatron chat endpoint writes (prefix_splice_suffix_token_ids,
prefix_splice_boundary_token_id). The key strings are spelled out here rather
than imported so this module stays importable in the vLLM worker and
finalizer environments; a test asserts parity with Megatron's constants when
Megatron is importable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Declare the nemo_gym extra on MegatronPolicyWorker in actor_environments.py
  (the venv source of truth) instead of swapping ACTOR_ENVIRONMENT_REGISTRY at
  runtime, which the prebuilt container venv ignored; drop PY_EXECUTABLES.MCORE_GYM
  and the ModuleNotFoundError string-match remediation.
- Gate backend=megatron token capture at setup on the MInf capture hook
  protocols (RequestPayloadStager / RequestPromptPreparer) from Megatron-LM
  PR #7015, failing with a NotImplementedError that names the dependency while
  the Megatron-Bridge pin predates it.
- Point the Gym submodule at Gym PR NVIDIA-NeMo#2823 (fc08bf19), which is reachable from
  NVIDIA-NeMo/Gym, fast-forwards from Gym main, and nests ng_capture under
  request_metadata as the Megatron endpoint requires.
- Skip test_prefix_splice_keys_match_megatron_constants when the pinned
  megatron-core lacks the constants, and importorskip megatron.core in the
  Megatron hosting test, so the Nemo_Gym shard skips instead of aborting.
- Restore the ++ Hydra override for checkpointing.save_data_plane in the
  streaming recovery script (the key is absent from the Gym config chain).
- Document the Megatron-LM #7015 dependency in the design doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Add design-docs/token-capture-ledger.md to the docs toctree and drop
  links to guides that do not exist yet (Sphinx treats both as errors).
- Read vllm_cfg / mcore_generation_config through a dict cast in the
  token-capture validation so pyrefly does not reject the TypedDict keys.
- Apply ruff formatting to megatron_worker.py and test_checkpointing.py.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…anup

Stamp a Megatron request that spans a refit with its admission (oldest)
policy epoch instead of failing capture and masking the whole rollout.
This matches vLLM, which freezes the version at begin_call, and the
finalizer's min-over-calls group tag; spans are counted
(epoch_span_count) and logged at WARNING.

Also fold in the mechanical review items: drop the undefined
rollout_max_attempts_to_avoid_lp_nan knob, remove the dead KeyError
branch in receipt parsing, route fetch_for_finalization through
TQStagingStore, collapse the duplicate prefix predicate, call
set_generation_epoch directly, fix import order, revert diff churn, fix
the FIFO docstring and the manifest control route in the design doc,
and remove the unused logical_request_id test parameter.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…dapter

TQMegatronTokenStager now hands the offloaded payload to
RolloutTokenCapture.complete_call_from_response with Gym's
MegatronCaptureAdapter instead of extracting fields by hand. Malformed
payloads poison the call with capture_failed coordinates, which Gym
records as worker_capture_failed (matching vLLM) rather than
worker_response_missing_commit_coordinates.

Requires the Gym-side adapter from NVIDIA-NeMo/Gym#2823.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- add parent_chain_hash to the inline-prefix token_in admission
- drop logical_request_id from _manifest_record (no such CallRecord field)

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- reword _assemble_receipt docstring so the named failure reasons are
  examples, not an exhaustive list; add unresolved_parent and the
  capture_failed fallback for reason-less rows
- terminal_selection is left unchanged: the pinned Gym RolloutReceipt
  (fc08bf19) types it Literal["declared","response_id","content",
  "heuristic"] with no None, so the invalid_manifest_row path cannot
  report "no heuristic ran" without a Gym-side schema change

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rics code

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wiring

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…fload_params

Megatron Inference renamed the opaque per-request dict it forwards to the
payload stager and prompt preparer (NVIDIA/Megatron-LM PR #7015); follow it
on the TQMegatronPromptPreparer / TQMegatronTokenStager side.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…_prefix_tokens

The Megatron chat endpoint now ships the chat-template render through the
last assistant message (template_prefix_token_ids) and the EOS id in
offload_params instead of a precomputed boundary (NVIDIA/Megatron-LM
PR #7015). TQMegatronPromptPreparer feeds those, plus the prefix it resolved
from the staging chain, to the same replace_prefix_tokens the vLLM worker
uses, so one splicing algorithm serves both backends.

replace_prefix_tokens gains a keyword-only eos_token_id override so callers
that hold only token ids can use it without a tokenizer; the boundary
failure message skips the detokenized reprs in that case. Existing callers
are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- bump Gym to 37dc751f (RolloutReceipt.terminal_selection is now Optional)
- _assemble_receipt no longer pre-labels parse failures as heuristic
- finalizer metric loop skips the None member of the Optional annotation

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Token capture on the Megatron backend was text-only: receipt rows trained
image-blind and every multi-turn VLM rollout failed at admission because the
preparer spliced the expanded previous turn into the compact new render.

- STAGING_FIELDS gains compact_token_ids_delta / compact_len; TQTokenSink pops
  Gym's compact delta out of extras into that column (like routed_experts) and
  TQTokenSource.fetch_prefix_chains returns both token spaces (compact falls
  back to expanded for text calls). ChainPrefixCache caches both.
- TQMegatronPromptPreparer splices with the compact chain, hands Gym the
  expanded chain to verify the engine prompt against, and records
  offload_params["ng_capture_minf"]["compact_prev_len"] for the stager.
- TQMegatronTokenStager derives the digest-covered media geometry from MInf's
  media_tensors for Gym's adapter, then stages the tensors themselves as extra
  columns on the call row (MEDIA_STAGING_FIELDS, a second put onto the same
  staging key once the token row is durable).
- RolloutReassembler reads the terminal call's media columns, checks them
  against the staged geometry (media_columns_missing / media_mismatch), and
  publishes the engine's packed patches as pixel_values with imgs_sizes and
  num_frames; the Megatron-Bridge Omni model accepts that layout directly.
- Setup registers the media columns for multimodal runs and rejects capture
  with a multimodal policy on the vLLM backend or with
  grpo.deduplicate_multimodal_data. The ledger design doc documents the flow.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the Documentation Improvements or additions to documentation label Sep 18, 2026
@lauradang

Copy link
Copy Markdown
Contributor Author

Superseded by lauradang#4, which is based on #4129's branch (laurad/minf-rollout-checkpointing) so the diff shows only the new commit.

@lauradang lauradang closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant