Conversation
Signed-off-by: Laura Dang <laurad@nvidia.com>
…very.sh Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <lauradang.2000@gmail.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
… cache Extract the vLLM worker's _fetch_chain_prefix and _resolve_admission_prefix into tq_token_sink.py as ChainPrefixCache and resolve_admission_prefix, and make the worker methods one-line delegates. TQMegatronPromptPreparer now resolves staging chains through the same pair, so both backends share one cached TQ read (256 entries keyed by the chain's last staging key, deepest cached key bounds the fetch to the uncached suffix). The preparer reads the splice boundary from the request-metadata keys the Megatron chat endpoint writes (prefix_splice_suffix_token_ids, prefix_splice_boundary_token_id). The key strings are spelled out here rather than imported so this module stays importable in the vLLM worker and finalizer environments; a test asserts parity with Megatron's constants when Megatron is importable. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
- Declare the nemo_gym extra on MegatronPolicyWorker in actor_environments.py (the venv source of truth) instead of swapping ACTOR_ENVIRONMENT_REGISTRY at runtime, which the prebuilt container venv ignored; drop PY_EXECUTABLES.MCORE_GYM and the ModuleNotFoundError string-match remediation. - Gate backend=megatron token capture at setup on the MInf capture hook protocols (RequestPayloadStager / RequestPromptPreparer) from Megatron-LM PR #7015, failing with a NotImplementedError that names the dependency while the Megatron-Bridge pin predates it. - Point the Gym submodule at Gym PR NVIDIA-NeMo#2823 (fc08bf19), which is reachable from NVIDIA-NeMo/Gym, fast-forwards from Gym main, and nests ng_capture under request_metadata as the Megatron endpoint requires. - Skip test_prefix_splice_keys_match_megatron_constants when the pinned megatron-core lacks the constants, and importorskip megatron.core in the Megatron hosting test, so the Nemo_Gym shard skips instead of aborting. - Restore the ++ Hydra override for checkpointing.save_data_plane in the streaming recovery script (the key is absent from the Gym config chain). - Document the Megatron-LM #7015 dependency in the design doc. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
- Add design-docs/token-capture-ledger.md to the docs toctree and drop links to guides that do not exist yet (Sphinx treats both as errors). - Read vllm_cfg / mcore_generation_config through a dict cast in the token-capture validation so pyrefly does not reject the TypedDict keys. - Apply ruff formatting to megatron_worker.py and test_checkpointing.py. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
…anup Stamp a Megatron request that spans a refit with its admission (oldest) policy epoch instead of failing capture and masking the whole rollout. This matches vLLM, which freezes the version at begin_call, and the finalizer's min-over-calls group tag; spans are counted (epoch_span_count) and logged at WARNING. Also fold in the mechanical review items: drop the undefined rollout_max_attempts_to_avoid_lp_nan knob, remove the dead KeyError branch in receipt parsing, route fetch_for_finalization through TQStagingStore, collapse the duplicate prefix predicate, call set_generation_epoch directly, fix import order, revert diff churn, fix the FIFO docstring and the manifest control route in the design doc, and remove the unused logical_request_id test parameter. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
…dapter TQMegatronTokenStager now hands the offloaded payload to RolloutTokenCapture.complete_call_from_response with Gym's MegatronCaptureAdapter instead of extracting fields by hand. Malformed payloads poison the call with capture_failed coordinates, which Gym records as worker_capture_failed (matching vLLM) rather than worker_response_missing_commit_coordinates. Requires the Gym-side adapter from NVIDIA-NeMo/Gym#2823. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
- add parent_chain_hash to the inline-prefix token_in admission - drop logical_request_id from _manifest_record (no such CallRecord field) Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- reword _assemble_receipt docstring so the named failure reasons are examples, not an exhaustive list; add unresolved_parent and the capture_failed fallback for reason-less rows - terminal_selection is left unchanged: the pinned Gym RolloutReceipt (fc08bf19) types it Literal["declared","response_id","content", "heuristic"] with no None, so the invalid_manifest_row path cannot report "no heuristic ran" without a Gym-side schema change Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rics code Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wiring Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…fload_params Megatron Inference renamed the opaque per-request dict it forwards to the payload stager and prompt preparer (NVIDIA/Megatron-LM PR #7015); follow it on the TQMegatronPromptPreparer / TQMegatronTokenStager side. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
…_prefix_tokens The Megatron chat endpoint now ships the chat-template render through the last assistant message (template_prefix_token_ids) and the EOS id in offload_params instead of a precomputed boundary (NVIDIA/Megatron-LM PR #7015). TQMegatronPromptPreparer feeds those, plus the prefix it resolved from the staging chain, to the same replace_prefix_tokens the vLLM worker uses, so one splicing algorithm serves both backends. replace_prefix_tokens gains a keyword-only eos_token_id override so callers that hold only token ids can use it without a tokenizer; the boundary failure message skips the detokenized reprs in that case. Existing callers are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Laura Dang <laurad@nvidia.com>
Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- bump Gym to 37dc751f (RolloutReceipt.terminal_selection is now Optional) - _assemble_receipt no longer pre-labels parse failures as heuristic - finalizer metric loop skips the None member of the Optional annotation Signed-off-by: Laura Dang <laurad@nvidia.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Laura Dang <laurad@nvidia.com>
Token capture on the Megatron backend was text-only: receipt rows trained image-blind and every multi-turn VLM rollout failed at admission because the preparer spliced the expanded previous turn into the compact new render. - STAGING_FIELDS gains compact_token_ids_delta / compact_len; TQTokenSink pops Gym's compact delta out of extras into that column (like routed_experts) and TQTokenSource.fetch_prefix_chains returns both token spaces (compact falls back to expanded for text calls). ChainPrefixCache caches both. - TQMegatronPromptPreparer splices with the compact chain, hands Gym the expanded chain to verify the engine prompt against, and records offload_params["ng_capture_minf"]["compact_prev_len"] for the stager. - TQMegatronTokenStager derives the digest-covered media geometry from MInf's media_tensors for Gym's adapter, then stages the tensors themselves as extra columns on the call row (MEDIA_STAGING_FIELDS, a second put onto the same staging key once the token row is durable). - RolloutReassembler reads the terminal call's media columns, checks them against the staged geometry (media_columns_missing / media_mismatch), and publishes the engine's packed patches as pixel_values with imgs_sizes and num_frames; the Megatron-Bridge Omni model accepts that layout directly. - Setup registers the media columns for multimodal runs and rejects capture with a multimodal policy on the vLLM backend or with grpo.deduplicate_multimodal_data. The ledger design doc documents the flow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Contributor
Author
|
Superseded by lauradang#4, which is based on #4129's branch (laurad/minf-rollout-checkpointing) so the diff shows only the new commit. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do ?
Stacked on #4129 (
laurad/minf-rollout-checkpointing); the diff shows that PR's commits until it merges. Only the top commit is new here.Token capture on the Megatron backend was text-only. Receipt rows carried no
pixel_values, so a VLM capture run trained image-blind, and every multi-turn VLM rollout failed at admission because the preparer spliced the expanded previous turn into the compact new render and Megatron expanded it twice.STAGING_FIELDSgainscompact_token_ids_delta/compact_len.TQTokenSinkpops Gym's compact delta out of extras into that column (likerouted_experts);TQTokenSource.fetch_prefix_chainsreturns both token spaces, with compact falling back to expanded for text calls.TQMegatronPromptPreparersplices with the compact chain, hands Gym the expanded chain to verify the engine prompt against, and recordsoffload_params["ng_capture_minf"]["compact_prev_len"]for the stager.TQMegatronTokenStagerderives the digest-covered media geometry from MInf'smedia_tensorsfor Gym's adapter, then stages the tensors themselves as extra columns on the same staging key (MEDIA_STAGING_FIELDS), the way routed experts ride the row.RolloutReassemblerreads the terminal call's media columns, checks them against the staged geometry (media_columns_missing/media_mismatch), and publishes the engine's packed patches aspixel_valueswithimgs_sizesandnum_frames. The Megatron-Bridge Omni model accepts the[total_patches, C*P*P]layout directly, so training projects exactly the pixels the policy generated against.grpo.deduplicate_multimodal_data.docs/design-docs/token-capture-ledger.mdgains a "Multimodal rollouts" section.Dependencies
compact_prompt_token_ids,media_tensorsonOffloadedRequestPayload), to be pinned through Megatron-Bridge.staging/media.py, Megatron adapter extras). Submodule anduv.lockre-pin pending, as with feat(rollout): integrate MInf ledger capture with checkpointing #4129.Testing
tests/unit/data_plane/test_tq_token_sink.pyandtest_rollout_reassembler.pyagainst the in-process TransferQueue simple backend: 56 passedtests/unit/single_controller/test_setup.py -k "multimodal or megatron_token_capture": 7 passedBefore your PR is "Ready for review"
🤖 Generated with Claude Code