Skip to content

feat(token-capture): carry media through Megatron capture rollouts - #4

Closed
lauradang wants to merge 96 commits into
pranav/vlm_capture_supportfrom
laurad/pr4129-multimodal-capture
Closed

lauradang wants to merge 96 commits into
pranav/vlm_capture_supportfrom
laurad/pr4129-multimodal-capture

Conversation

@lauradang

@lauradang lauradang commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

What does this PR do ?

Stacked on NVIDIA-NeMo#4129 (laurad/minf-rollout-checkpointing), which is this PR's base branch.

Token capture on the Megatron backend was text-only. Receipt rows carried no pixel_values, so a VLM capture run trained image-blind, and every multi-turn VLM rollout failed at admission because the preparer spliced the expanded previous turn into the compact new render and Megatron expanded it twice.

  • Compact chains. STAGING_FIELDS gains compact_token_ids_delta / compact_len. TQTokenSink pops Gym's compact delta out of extras into that column (like routed_experts); TQTokenSource.fetch_prefix_chains returns both token spaces, with compact falling back to expanded for text calls.
  • Preparer. TQMegatronPromptPreparer splices with the compact chain, hands Gym the expanded chain to verify the engine prompt against, and records offload_params["ng_capture_minf"]["compact_prev_len"] for the stager.
  • Media on the call row, as per-call deltas. Every chat request carries the whole conversation, so the engine hands the stager pixels for every image in the prompt. The preparer records how many media items the parent chain already staged (media_prev_count, next to compact_prev_len), the stager slices MInf's media_tensors at that boundary, derives the digest-covered geometry for Gym's adapter from the remainder, and stages those tensors as extra columns on the same staging key (MEDIA_STAGING_FIELDS), the way routed experts ride the row. Each row holds only the media new to that call, like the token columns.
  • Finalizer. RolloutReassembler walks the terminal chain, checks each call's media columns against that call's staged geometry (media_columns_missing / media_mismatch), concatenates the deltas in chain order as it does for tokens, and publishes the engine's packed patches as pixel_values with imgs_sizes and num_frames. The Megatron-Bridge Omni model accepts the [total_patches, C*P*P] layout directly, so training projects exactly the pixels the policy generated against.
  • Setup. Registers the media columns for multimodal runs; rejects capture with a multimodal policy on the vLLM backend (that path stages the pre-processor prompt and carries no media) and with grpo.deduplicate_multimodal_data.

docs/design-docs/token-capture-ledger.md gains a "Multimodal rollouts" section.

Worked example

A two-turn VLM rollout under capture, with toy values: the media placeholder is token 99, EOS is 2, one 4x4 image expands to three media tokens and to four packed patches of 12 features. Turn 1 sends text plus one image. Turn 2 sends the whole history plus a second image; the chat template re-tokenizes the turn-1 answer as 13 although the model generated 12, which is what the splice corrects. Each box is a process; the label says whose code runs in it.

sequenceDiagram
    box Gym server (Gym code)
        participant G as Agent + capture middleware<br/>+ MegatronWorkerCaptureHandler
    end
    box Megatron HTTP endpoint, in RL's Megatron worker (Megatron code)
        participant EP as chat_completions.py
    end
    box Engine TP0/PP0 (Megatron loop, RL+Gym hooks)
        participant P as TQMegatronPromptPreparer
        participant S as TQMegatronTokenStager<br/>+ Gym adapter
    end
    box Engine, all ranks (Megatron)
        participant E as add_request / _build_vlm_request
    end
    participant TQ as TransferQueue
    participant L as Gym ledger
    box RL actors (RL code)
        participant F as RolloutReassembler
        participant T as Trainer
    end

    Note over G,L: Turn 1 - text + image1
    G->>EP: chat req, offload_params.ng_capture = {c1, text, prev_len 0}
    EP->>P: compact [80,99,81] + image bytes (via DP coordinator)
    P->>E: text mode, pass through, broadcast
    E->>E: decode img1 to imgs[1,4,12], expand to [80,99,99,99,81], gen [12,2]
    E->>S: payload {expanded, gen, compact (new), media_tensors (new)}
    S->>TQ: r0/c1 - tokens, compact col [80,99,81,12,2], extras {imgs_sizes}, media cols
    S->>EP: ng_commit_coords {r0/c1, cum_len 7} via engine reply
    EP->>G: "a cat" + coords
    G->>L: CallRecord c1

    Note over G,L: Turn 2 - history + image2
    G->>L: resolve_parent finds c1
    G->>EP: ng_capture = {c2, token_in, prev_len 7, staging_chain [r0/c1]}
    EP->>P: compact [80,99,81,13,2,20,99,21], template_prefix [80,99,81,13,2], eos 2
    P->>TQ: fetch_prefix_chains([r0/c1])
    TQ->>P: expanded [80,99,99,99,81,12,2] / compact [80,99,81,12,2]
    P->>E: splice compact to [80,99,81,12,2,20,99,21], required_prefix = expanded, compact_prev_len = 5
    E->>E: 2 placeholders = 2 images OK, expand to [80,99,99,99,81,12,2,20,99,99,99,21], gen [30,2]
    E->>S: payload
    S->>S: Gym checks prompt[:7] == expanded chain, delta [20,99,99,99,21,30,2], compact [20,99,21,30,2]
    S->>TQ: r0/c2 tokens + compact + extras + media cols (image 2 only, media_prev_count 1)
    S->>EP: coords {r0/c2, prev_len 7, cum_len 14}
    EP->>G: answer + coords
    G->>L: CallRecord c2 (parent c1)

    Note over F,T: Rollout end
    F->>L: manifest to receipt {terminal c2, [c1, c2]}
    F->>TQ: fetch rows for 14 expanded tokens, fetch_media for r0/c1 and r0/c2, concat to imgs[1,8,12]
    F->>F: imgs_sizes cols == extras OK, pixel_values[8,12]
    F->>TQ: publish canonical row, clear r0/c1 and r0/c2
    T->>TQ: read input_ids + pixel_values, six 99s match six embeddings
Loading

Where the old code failed: at turn 2 the preparer spliced the expanded chain, so the engine saw four 99s for two images and rejected the request on the placeholder count. Splicing in compact space fixes that, and Gym's existing prefix check on the expanded prompt (prompt[:7]) verifies that re-expanding image 1 reproduced the same tokens.

Dependencies

Testing

  • tests/unit/data_plane/test_tq_token_sink.py and test_rollout_reassembler.py against the in-process TransferQueue simple backend: 63 passed
  • tests/unit/single_controller/test_setup.py -k "multimodal or megatron_token_capture": 7 passed

End-to-end Nemotron Omni video run

Completed a 42-step, 8-node × 4-GPU VSTAT GRPO run in 19m14s with this PR stack:

  • RL: 79fa6856b5e40a456c8d932f434a490c9682e296
  • Gym: fafecf7c71ef63d027bb2197895e93c9effced4a
  • Megatron-LM: 0aa6cce27f75c2a4cf733b8eeb0e8180dc451707
  • Megatron-Bridge: 8e4ccc34b67a79bca73cb76242479f9090856ff4

Run: PR stack (jtg9a7rp) · cye baseline (o4ed2p8n) · W&B overlay

The reward curves track closely: Pearson correlation 0.9169, mean absolute error 0.0486; mean reward over all 42 steps was 0.1692 for this PR stack vs. 0.1592 for the baseline. Steps 23–42 averaged 0.2229 vs. 0.2281.

VSTAT reward overlay using nemo_rl/step

The chart uses nemo_rl/step on the x-axis (not W&B's synthetic STEP). Duplicate history entries from continuation were reduced to the final value for each nemo_rl/step.

Important effective configuration:

policy:
  model_name: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
  is_vlm: true
  megatron_cfg:
    freeze_vision_model: false
    freeze_vision_projection: false
  generation:
    backend: megatron
    mcore_generation_config:
      megatron_inference_wrapper: megatron.core.inference.model_inference_wrappers.multimodal.nemotron_omni_inference_wrapper.NemotronOmniInferenceWrapper
      parsers: [nemotron-v3-reasoning, qwen3-coder-tool]
      video_num_frames: 16
      video_temporal_patch_size: 2
      video_target_num_patches: 1024

data:
  train:
    data_path: /opt/nemo-rl/workspace/datasets/vstat-8n4g/train-gym.jsonl
  default:
    dataset_name: NemoGymDataset
    processor: nemo_gym_data_processor
    num_frames: 16
    video_sampling_style: nemotron_vl
    video_temporal_patch_size: 2
    video_target_num_patches: 1024
    video_maintain_aspect_ratio: true

grpo:
  deduplicate_multimodal_data: false

data_plane:
  enabled: true
  impl: transfer_queue
  backend: simple

token_capture:
  enabled: true

checkpointing:
  enabled: true
  save_period: 10
  save_data_plane: true

The dataset rows contain real local video paths. The step-40 native TQ checkpoint contains pixel_values, imgs_sizes, and num_frames (plus their row-shape metadata) in replay_buffer_metadata.pt, and the same media columns in the TQ storage-unit pickle. A continuation read those columns from TQ and reached Nemotron Omni's _encode_images(images, ..., num_frames) path during policy logprob computation.

🤖 Generated with Claude Code

lauradang and others added 26 commits September 14, 2026 13:09
Signed-off-by: Laura Dang <laurad@nvidia.com>
…very.sh

Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… cache

Extract the vLLM worker's _fetch_chain_prefix and _resolve_admission_prefix
into tq_token_sink.py as ChainPrefixCache and resolve_admission_prefix, and
make the worker methods one-line delegates. TQMegatronPromptPreparer now
resolves staging chains through the same pair, so both backends share one
cached TQ read (256 entries keyed by the chain's last staging key, deepest
cached key bounds the fetch to the uncached suffix).

The preparer reads the splice boundary from the request-metadata keys the
Megatron chat endpoint writes (prefix_splice_suffix_token_ids,
prefix_splice_boundary_token_id). The key strings are spelled out here rather
than imported so this module stays importable in the vLLM worker and
finalizer environments; a test asserts parity with Megatron's constants when
Megatron is importable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Declare the nemo_gym extra on MegatronPolicyWorker in actor_environments.py
  (the venv source of truth) instead of swapping ACTOR_ENVIRONMENT_REGISTRY at
  runtime, which the prebuilt container venv ignored; drop PY_EXECUTABLES.MCORE_GYM
  and the ModuleNotFoundError string-match remediation.
- Gate backend=megatron token capture at setup on the MInf capture hook
  protocols (RequestPayloadStager / RequestPromptPreparer) from Megatron-LM
  PR #7015, failing with a NotImplementedError that names the dependency while
  the Megatron-Bridge pin predates it.
- Point the Gym submodule at Gym PR NVIDIA-NeMo#2823 (fc08bf19), which is reachable from
  NVIDIA-NeMo/Gym, fast-forwards from Gym main, and nests ng_capture under
  request_metadata as the Megatron endpoint requires.
- Skip test_prefix_splice_keys_match_megatron_constants when the pinned
  megatron-core lacks the constants, and importorskip megatron.core in the
  Megatron hosting test, so the Nemo_Gym shard skips instead of aborting.
- Restore the ++ Hydra override for checkpointing.save_data_plane in the
  streaming recovery script (the key is absent from the Gym config chain).
- Document the Megatron-LM #7015 dependency in the design doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Add design-docs/token-capture-ledger.md to the docs toctree and drop
  links to guides that do not exist yet (Sphinx treats both as errors).
- Read vllm_cfg / mcore_generation_config through a dict cast in the
  token-capture validation so pyrefly does not reject the TypedDict keys.
- Apply ruff formatting to megatron_worker.py and test_checkpointing.py.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…anup

Stamp a Megatron request that spans a refit with its admission (oldest)
policy epoch instead of failing capture and masking the whole rollout.
This matches vLLM, which freezes the version at begin_call, and the
finalizer's min-over-calls group tag; spans are counted
(epoch_span_count) and logged at WARNING.

Also fold in the mechanical review items: drop the undefined
rollout_max_attempts_to_avoid_lp_nan knob, remove the dead KeyError
branch in receipt parsing, route fetch_for_finalization through
TQStagingStore, collapse the duplicate prefix predicate, call
set_generation_epoch directly, fix import order, revert diff churn, fix
the FIFO docstring and the manifest control route in the design doc,
and remove the unused logical_request_id test parameter.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…dapter

TQMegatronTokenStager now hands the offloaded payload to
RolloutTokenCapture.complete_call_from_response with Gym's
MegatronCaptureAdapter instead of extracting fields by hand. Malformed
payloads poison the call with capture_failed coordinates, which Gym
records as worker_capture_failed (matching vLLM) rather than
worker_response_missing_commit_coordinates.

Requires the Gym-side adapter from NVIDIA-NeMo/Gym#2823.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- add parent_chain_hash to the inline-prefix token_in admission
- drop logical_request_id from _manifest_record (no such CallRecord field)

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- reword _assemble_receipt docstring so the named failure reasons are
  examples, not an exhaustive list; add unresolved_parent and the
  capture_failed fallback for reason-less rows
- terminal_selection is left unchanged: the pinned Gym RolloutReceipt
  (fc08bf19) types it Literal["declared","response_id","content",
  "heuristic"] with no None, so the invalid_manifest_row path cannot
  report "no heuristic ran" without a Gym-side schema change

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rics code

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wiring

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…fload_params

Megatron Inference renamed the opaque per-request dict it forwards to the
payload stager and prompt preparer (NVIDIA/Megatron-LM PR #7015); follow it
on the TQMegatronPromptPreparer / TQMegatronTokenStager side.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…_prefix_tokens

The Megatron chat endpoint now ships the chat-template render through the
last assistant message (template_prefix_token_ids) and the EOS id in
offload_params instead of a precomputed boundary (NVIDIA/Megatron-LM
PR #7015). TQMegatronPromptPreparer feeds those, plus the prefix it resolved
from the staging chain, to the same replace_prefix_tokens the vLLM worker
uses, so one splicing algorithm serves both backends.

replace_prefix_tokens gains a keyword-only eos_token_id override so callers
that hold only token ids can use it without a tokenizer; the boundary
failure message skips the detokenized reprs in that case. Existing callers
are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- bump Gym to 37dc751f (RolloutReceipt.terminal_selection is now Optional)
- _assemble_receipt no longer pre-labels parse failures as heuristic
- finalizer metric loop skips the None member of the Optional annotation

Signed-off-by: Laura Dang <laurad@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Keep the Gym submodule at 37dc751f (Gym NVIDIA-NeMo#2823 head); main's pin 267305e2 is an ancestor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Token capture on the Megatron backend was text-only: receipt rows trained
image-blind and every multi-turn VLM rollout failed at admission because the
preparer spliced the expanded previous turn into the compact new render.

- STAGING_FIELDS gains compact_token_ids_delta / compact_len; TQTokenSink pops
  Gym's compact delta out of extras into that column (like routed_experts) and
  TQTokenSource.fetch_prefix_chains returns both token spaces (compact falls
  back to expanded for text calls). ChainPrefixCache caches both.
- TQMegatronPromptPreparer splices with the compact chain, hands Gym the
  expanded chain to verify the engine prompt against, and records
  offload_params["ng_capture_minf"]["compact_prev_len"] for the stager.
- TQMegatronTokenStager derives the digest-covered media geometry from MInf's
  media_tensors for Gym's adapter, then stages the tensors themselves as extra
  columns on the call row (MEDIA_STAGING_FIELDS, a second put onto the same
  staging key once the token row is durable).
- RolloutReassembler reads the terminal call's media columns, checks them
  against the staged geometry (media_columns_missing / media_mismatch), and
  publishes the engine's packed patches as pixel_values with imgs_sizes and
  num_frames; the Megatron-Bridge Omni model accepts that layout directly.
- Setup registers the media columns for multimodal runs and rejects capture
  with a multimodal policy on the vLLM backend or with
  grpo.deduplicate_multimodal_data. The ledger design doc documents the flow.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 18, 2026
Every chat request carries the whole conversation, so the engine hands the
stager pixels for every image in the prompt and each call row stored all of
them: a T-turn rollout kept a turn-1 image T times until cleanup. Media
columns now follow the same rule as the token columns and hold only what is
new to the call.

- The preparer counts the media items in the parent rows' staged geometry
  while resolving the chain and records media_prev_count next to
  compact_prev_len in offload_params["ng_capture_minf"].
- The stager slices the engine's media tensors at that boundary before Gym's
  adapter sees them (slice_media_tensors), so both the digest-covered geometry
  and the media columns describe only this call's media.
- The finalizer walks the terminal chain, verifies each call's media columns
  against its own geometry, and concatenates the deltas in order, as it does
  for tokens. The published pixel_values are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@lauradang lauradang self-assigned this Sep 18, 2026
lauradang and others added 5 commits September 23, 2026 12:36
…egatron stager poison path

- Defer the openai_server_utils import into TQMegatronPromptPreparer
  .prepare_prompt so importing tq_token_sink no longer pulls
  nemo_rl.models.generation (transformers, vLLM, TRT-LLM) into the
  finalizer and controller import paths.
- _MegatronCapturePayload.from_offloaded now raises TypeError for a
  non-mapping media_tensors (previously escaped slice_media_tensors as
  AttributeError) and for non-dict capture params (previously treated
  as empty), so both land on the capture_failed poison path.
- Document that TQTokenSource only computes media_count when built with
  capture_media=True, so the preparer's media_prev_count depends on the
  source and sink flags matching (both set in setup_token_capture).
- Delete the unused resolve_admission_prefix and ChainPrefixCache.fetch
  wrappers; the vLLM worker has its own equivalents and the Megatron
  preparer uses resolve_admission_prefix_chains directly.
- Tests: compact_token_ids_delta rejection in stage(), compact_len
  mismatch on read, non-mapping poison cases, media_tensors={} yields
  attachments=None, and drop a MagicMock-spec-only assertion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…ansport

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…e-capable wrapper; pin MInf pixel dtype

Refuse capture_media=True at setup_token_capture when the worker built its
engine without image preprocessing (a text-only wrapper never yields media
tensors), assert the setup test against the returned gen_handle, pin the
float32 premise behind MINF_MEDIA_PIXEL_DTYPE with an mcore test, and run the
finalizer's attached-media case over both backends' pixel dtypes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…a_tensors for the slicer

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…load fields

Fail at setup when the pinned Megatron-LM lacks OffloadedRequestPayload.media_tensors
instead of staging sentinels and dropping every group; mark the payload-field pin test
xfail(strict) until the Megatron-LM re-pin lands.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
@lauradang
lauradang changed the base branch from laurad/minf-rollout-checkpointing to pranav/vlm_capture_support September 23, 2026 21:21
cspades and others added 20 commits September 23, 2026 15:10
Replace the remaining bare "ng_capture" / "ng_commit_coords" literals in
TQMegatronPromptPreparer.prepare_prompt and TQMegatronTokenStager with
NG_CAPTURE_FIELD / NG_COMMIT_COORDS_FIELD exported by
nemo_gym.token_id_capture, so a Gym rename fails at import time instead
of silently declining every request.

Addresses NVIDIA-NeMo#4129 (comment)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
CI's lint job fails when a file type-checks clean but is absent from
pyrefly.toml's project-includes; the new Megatron capture module was
missing, which failed Lint check and skipped every downstream test job
on the previous PR head.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… traces (NVIDIA-NeMo#4052)

Signed-off-by: Raj Singh <rajsin@nvidia.com>
Signed-off-by: Nolen Liang <nliang@nvidia.com>
Signed-off-by: kajalj <kajalj@nvidia.com>
Signed-off-by: Cory Ye <cye@nvidia.com>
Co-authored-by: Nolen Liang <nliang@nvidia.com>
Co-authored-by: Cory Ye <cye@nvidia.com>
Conflict resolutions:
- Megatron-Bridge submodule: take main's 1f8873bb (NVIDIA-NeMo#4139); it already contains
  the PR's 4d472695 gather_output fix.
- community_import.py / test_community_import.py: take main, which removed the
  _prefer_nvrx_for_dist_ckpt_save shim the PR had extended with a signature check.
- megatron_policy_worker.py: take main's unconditional FileSystemWriterAsync
  import and direct cleanup_tensor_caches() call.
- L1_Functional_Tests_SingleController: main split the script into _1/_2/_3;
  the PR's Megatron sibling-recovery variant now lives in _3 next to its vLLM twin.
- pyproject.toml: keep the PR's comment on the flashinfer index URL spelling.
- uv.lock: relocked with uv 0.11.28 (Dockerfile pin); result matches main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…dget

- test_backend_capture_glue_reproduces_the_gym_worked_example compared coords
  against every CallRecord field except mode/response_id, but CallRecord also
  carries attribution-only fields (admitted_at, output_fingerprint,
  continuation_fingerprint, fingerprint_version) that CommitCoords never has,
  so the lookup raised KeyError on whichever one the set yielded first.
  Compare on the fields the two models share.
- The new token-capture nightly cost 48 GPU-hours and pushed the suite from
  4698 to 4746, over the 4720 cap asserted by
  test_nightly_compute_stays_below_4720_hours. Cut it to 3 steps / 82 min
  (21 GPU-hours) so the suite lands at 4719.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…IA-NeMo#4129)

Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Terry Kong <terrycurtiskong@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…A-NeMo#4260)

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Extend NeMo-Gym token capture to Omni dynamic-resolution images and native
video for async vLLM generation with a Megatron learner. The vLLM worker
snapshots the processed media tensors the engine ran on, stages them in the
same TransferQueue write as each call's token delta, and the finalizer
reassembles the rollout's media chain into pixel_values / imgs_sizes /
num_frames for the learner. Text-only runs are unaffected.

Rebased onto main after the MInf ledger-capture and vLLM 0.29 landings:

- reuse main's shared ChainPrefixCache / resolve_admission_prefix in the
  vLLM worker; the media path keeps its own handle on the same TQTokenSource
- declare capture_media on GenerationInterface.setup_token_capture and reject
  it explicitly in MegatronGeneration; setup fails at validation time when
  media capture is requested with a non-vLLM generation backend
- keep main's Optional-aware terminal_selection metric derivation
- Gym pin 17abf2dd (includes main's pin plus the attachment staging contract);
  uv.lock reflects Gym dropping httpx-aiohttp

Paired Gym change: NVIDIA-NeMo/Gym#3513.

Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main added an eos_token_id override to replace_prefix_tokens for callers
without a tokenizer (the Megatron prompt preparer). The media-capture split
into splice_prefix_tokens dropped the parameter, so the splice body read an
undefined name and every non-trivial replace_prefix_tokens call raised.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
…M 0.29

The fake capture workers gain the ChainPrefixCache the rebased worker
resolves admissions through, and the real-processor test disables vLLM
0.29's device-side normalization on its bare MultiModalConfig: real engines
switch it off for the Omni model (no supports_mm_device_do_normalize), but a
bare config never runs that resolution and the injected do_normalize kwarg
breaks the processor constructor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
… via signature

main already expects the per-agent mask_sample metrics in test_rollouts, so
the branch's copy became a repeated dict literal after the rebase (ruff F601);
take main's file. MultiModalConfig is a dataclass in vLLM 0.29, so detect the
mm_device_do_normalize field through inspect.signature instead of
pydantic's model_fields.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
…or test

NanoNemotronVLMultiModalProcessor now runs the HF processor on dummy text
(_get_hf_processor_text), so a request's own text no longer competes with
its images for tiles. Only adding an image under a tight budget re-tiles a
retained one; update the expectation matrix and the guide's caveat.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Pranav Prashant Thombre <pthombre@nvidia.com>
…l-capture

Brings in the final NVIDIA-NeMo#4129 branch (squashed upstream as e073dab) so the
upstream vlm_capture_support merge can use the squash as its base.

Conflict resolution:
- NVIDIA-NeMo#4129 moved TQMegatronPromptPreparer / TQMegatronTokenStager to
  nemo_rl/models/generation/megatron/token_capture.py; port the media
  capture changes (PrefixChains / resolve_admission_prefix_chains, compact
  splice, _MegatronCapturePayload, media attachments) onto that module,
  keeping its NG_*_FIELD constants, megatron-core PREFIX_* fields and
  RequestPayloadStageResult.
- NVIDIA-NeMo#4129 deleted the driver-side _require_minf_capture_hooks gate; keep only
  _require_minf_media_payload_fields for media runs.
- Keep the multimodal Megatron stager tests alongside NVIDIA-NeMo#4129's backend
  parity test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Re-resolved the NVIDIA-NeMo#4129 squash-merge conflicts with the squash (e073dab)
as the merge-file base; its tree matches the fork head merged in the
previous commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
…/pr4129-multimodal-capture

Retargets this branch at the rebased NVIDIA-NeMo#4191 head. NVIDIA-NeMo#4191 was rewritten on
top of main 4f4f351, so its old commits (up to 90a25ca) that this
branch already carried were re-resolved against a synthetic base of main
plus those old changes; only the new NVIDIA-NeMo#4191 changes are real conflicts.

Resolution:
- vLLM worker/generation, their hosting tests, test_captured_media and the
  Omni video functional test take NVIDIA-NeMo#4191 as-is (no PR 4 changes there).
- tq_token_sink: one ChainPrefixCache serving both backends: the flat
  fetch()/resolve_admission_prefix() the vLLM worker uses plus
  fetch_chains()/resolve_admission_prefix_chains() for the Megatron
  preparer; keep the compact-delta columns and add NVIDIA-NeMo#4191's media metadata
  digest (MEDIA_METADATA_FIELDS) on the same put.
- Drop NVIDIA-NeMo#4191's vLLM-only media guards in setup_single_controller and
  MegatronGeneration.setup_token_capture; this PR enables Megatron media
  capture (gated on _require_minf_media_payload_fields).
- Restore NVIDIA-NeMo#4129's receipt check (the empty_manifest rejection had come
  back through the earlier NVIDIA-NeMo#4191 merge).
- Pin Gym to lauradang/Gym laurad/pr2823-multimodal-capture a2ef85b0e,
  which merges NVIDIA-NeMo#3513's squash (17abf2ddc) into the Megatron adapter branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
@lauradang

Copy link
Copy Markdown
Owner Author

Superseded by NVIDIA-NeMo#4332, which targets the upstream pranav/vlm_capture_support (NVIDIA-NeMo#4191) directly. Same head branch; it now merges the rebased NVIDIA-NeMo#4191 head 98e6a8cea.

@lauradang lauradang closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.