Parent
Child of #3224 and #3024.
Problem
The Megatron inference integration currently expects prefix token data to be supplied from Gym/framework-side dispatch and does not have the same direct TransferQueue lineage traversal available to vLLM workers. A staged-chain-only recovery contract would therefore make checkpoint continuation backend-specific and unusable for this path.
Proposed implementation
- Implement the common generation-cut inventory and acknowledgement contract at the Megatron generation boundary.
- Persist Megatron-owned generation metadata needed to continue safely.
- Use the explicit inline-prefix recovery representation when the generation worker cannot dereference TQ coordinates itself.
- Have the framework integration materialize the staged prefix and validate its digest before dispatch.
- Keep inline and staged prefix variants mutually exclusive while sharing model-call and token-boundary identity.
Related work: Gym #2823, NeMo-RL #3869, and NVIDIA/Megatron-LM#7015.
Acceptance criteria
Parent
Child of #3224 and #3024.
Problem
The Megatron inference integration currently expects prefix token data to be supplied from Gym/framework-side dispatch and does not have the same direct TransferQueue lineage traversal available to vLLM workers. A staged-chain-only recovery contract would therefore make checkpoint continuation backend-specific and unusable for this path.
Proposed implementation
Related work: Gym #2823, NeMo-RL #3869, and NVIDIA/Megatron-LM#7015.
Acceptance criteria