Skip to content

Adding nemo template - #1086

Merged
Aydin-ab merged 5 commits into
mainfrom
template/nemotron-example
Sep 14, 2026
Merged

Aydin-ab merged 5 commits into
mainfrom
template/nemotron-example

Conversation

@xing-anyscale

@xing-anyscale xing-anyscale commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

workspace

Readme

notebook

I have run the template in workspace to make sure it's runnable.

@xing-anyscale
xing-anyscale force-pushed the template/nemotron-example branch from 754302c to 20d4629 Compare September 11, 2026 21:31
@xing-anyscale
xing-anyscale force-pushed the template/nemotron-example branch from 20d4629 to 926e725 Compare September 11, 2026 21:35
@xing-anyscale
xing-anyscale marked this pull request as ready for review September 11, 2026 21:39
mamba_ssm_cache_dtype="float32",
gpu_memory_utilization=0.9,
trust_remote_code=True,
enable_chunked_prefill=True,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is the default if I am not mistaken?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch — confirmed, dropping enable_chunked_prefill=True as SchedulerConfig.enable_chunked_prefill defaults to True in vLLM 0.25.1 (V1 enables chunked
prefill by default), so the line is a no-op.

Keeping the other three — none of them are defaults on 0.25.1:

  • gpu_memory_utilization=0.9 — default is 0.92, and 0.9 is what NVIDIA's card specifies
  • mamba_ssm_cache_dtype="float32" — default is "auto"
  • trust_remote_code=True — default is False, and required for the custom nemotron_h
    architecture

# async_scheduling=True,
# max_cudagraph_capture_size=128,
reasoning_parser="nemotron_v3",
# enable_auto_tool_choice=True, # uncomment for agents/tool use

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make these the default.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call. will do.

Comment on lines +51 to +52
# async_scheduling=True,
# max_cudagraph_capture_size=128,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

async scheduling should be the default on the latest ray serve llm image. These seem to be referencing an old vllm version, why?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed on both. let me update it.

)
),
# runtime_env=dict(env_vars={"HF_TOKEN": os.environ.get("HF_TOKEN")}),
engine_kwargs=dict(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

did you cross check this with the vllm recipes's page?

vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
--trust-remote-code
--kv-cache-dtype fp8
--tensor-parallel-size 8
--enable-auto-tool-choice
--tool-call-parser qwen3_xml
--reasoning-parser nemotron_v3
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'

https://recipes.vllm.ai/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16?hardware=h100&features=tool_calling,reasoning,spec_decoding&variant=fp8

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

there is spec decoding. It's tp8 and no ep here.

@xing-anyscale xing-anyscale Sep 11, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call — let me follow the recipe and update the config: TP=8, drop expert parallelism, addspeculative_config={"method": "mtp", "num_speculative_tokens": 3}. Taking all three together since that's the combination recipes validated.

Comment on lines +90 to +91
!pip install "ray[serve,llm]==2.57.0"
!pip install "vllm==0.25.1"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why are we on an old version here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i discussed this with Kunling as well. Since other templates in this deployment-serve-llm folder are all based on ray=2..57, we wanna keep the same version for this one. lmk if it makes sense. @kunling-anyscale

Comment thread templates/deployment-serve-llm/nemotron-3-super-120b/workspace_prod.yaml Outdated
@xing-anyscale
xing-anyscale force-pushed the template/nemotron-example branch from 609f4c8 to 08eb20e Compare September 12, 2026 00:23

@Aydin-ab Aydin-ab left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd like to make that troubleshooting more concise / nemotron-specific to make it more readable

Otherwise LGTM thank you for adding this 🙏

I believe ray docs should sync automatically but cc @elliot-barn

Comment thread templates/deployment-serve-llm/nemotron-3-super-120b/README.md Outdated
@Aydin-ab

Copy link
Copy Markdown
Contributor

I can do the bump to 2.58 once this merges

Aydin-ab and others added 2 commits September 14, 2026 12:34
8a8fd8b tightened the Troubleshooting section in README.md directly, which
desynced it from README.ipynb and failed the generate-readme hook. The notebook
is the generated source, so port the same text there; README.md is unchanged and
now matches nbconvert output.

Signed-off-by: Aydin Abiar <aydin@anyscale.com>
@Aydin-ab
Aydin-ab merged commit b19087d into main Sep 14, 2026
4 checks passed
@Aydin-ab
Aydin-ab deleted the template/nemotron-example branch September 14, 2026 20:00
Aydin-ab pushed a commit that referenced this pull request Sep 14, 2026
…#1087)

#1086 added the Nemotron-3-Super-120B template but didn't list it here,
so the index stops at gpt-oss.

Adding it surfaced a second problem: all seven existing tutorial links
point at
`serve/tutorials/deployment-serve-llm/content/<name>/README.html`, which
redirects to the Serve index. Ray publishes these under `_collections/`
— per `doc/source/serve/examples.yml`, the link is
`../_collections/serve/tutorials/deployment-serve-llm/<name>/README`.
Repointed all eight.

## Testing

`curl -sL` on each of the seven repointed links returns HTTP 200 at the
requested path, against a soft 404 before. The Nemotron link 404s until
Ray picks the template up in its next docs sync — the same lag every new
tutorial has.

Signed-off-by: svc-template-updater <282975472+svc-template-updater@users.noreply.github.com>
Co-authored-by: svc-template-updater <282975472+svc-template-updater@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants