Adding nemo template - #1086
Adding nemo template#1086
Conversation
754302c to
20d4629
Compare
20d4629 to
926e725
Compare
| mamba_ssm_cache_dtype="float32", | ||
| gpu_memory_utilization=0.9, | ||
| trust_remote_code=True, | ||
| enable_chunked_prefill=True, |
There was a problem hiding this comment.
this is the default if I am not mistaken?
There was a problem hiding this comment.
Nice catch — confirmed, dropping enable_chunked_prefill=True as SchedulerConfig.enable_chunked_prefill defaults to True in vLLM 0.25.1 (V1 enables chunked
prefill by default), so the line is a no-op.
Keeping the other three — none of them are defaults on 0.25.1:
gpu_memory_utilization=0.9— default is 0.92, and 0.9 is what NVIDIA's card specifiesmamba_ssm_cache_dtype="float32"— default is"auto"trust_remote_code=True— default isFalse, and required for the customnemotron_h
architecture
| # async_scheduling=True, | ||
| # max_cudagraph_capture_size=128, | ||
| reasoning_parser="nemotron_v3", | ||
| # enable_auto_tool_choice=True, # uncomment for agents/tool use |
There was a problem hiding this comment.
should we make these the default.
There was a problem hiding this comment.
Good call. will do.
| # async_scheduling=True, | ||
| # max_cudagraph_capture_size=128, |
There was a problem hiding this comment.
async scheduling should be the default on the latest ray serve llm image. These seem to be referencing an old vllm version, why?
There was a problem hiding this comment.
Agreed on both. let me update it.
| ) | ||
| ), | ||
| # runtime_env=dict(env_vars={"HF_TOKEN": os.environ.get("HF_TOKEN")}), | ||
| engine_kwargs=dict( |
There was a problem hiding this comment.
did you cross check this with the vllm recipes's page?
vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
--trust-remote-code
--kv-cache-dtype fp8
--tensor-parallel-size 8
--enable-auto-tool-choice
--tool-call-parser qwen3_xml
--reasoning-parser nemotron_v3
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
There was a problem hiding this comment.
there is spec decoding. It's tp8 and no ep here.
There was a problem hiding this comment.
Good call — let me follow the recipe and update the config: TP=8, drop expert parallelism, addspeculative_config={"method": "mtp", "num_speculative_tokens": 3}. Taking all three together since that's the combination recipes validated.
| !pip install "ray[serve,llm]==2.57.0" | ||
| !pip install "vllm==0.25.1" |
There was a problem hiding this comment.
why are we on an old version here?
There was a problem hiding this comment.
i discussed this with Kunling as well. Since other templates in this deployment-serve-llm folder are all based on ray=2..57, we wanna keep the same version for this one. lmk if it makes sense. @kunling-anyscale
609f4c8 to
08eb20e
Compare
08eb20e to
8ca95f0
Compare
Aydin-ab
left a comment
There was a problem hiding this comment.
I'd like to make that troubleshooting more concise / nemotron-specific to make it more readable
Otherwise LGTM thank you for adding this 🙏
I believe ray docs should sync automatically but cc @elliot-barn
|
I can do the bump to 2.58 once this merges |
8a8fd8b tightened the Troubleshooting section in README.md directly, which desynced it from README.ipynb and failed the generate-readme hook. The notebook is the generated source, so port the same text there; README.md is unchanged and now matches nbconvert output. Signed-off-by: Aydin Abiar <aydin@anyscale.com>
…#1087) #1086 added the Nemotron-3-Super-120B template but didn't list it here, so the index stops at gpt-oss. Adding it surfaced a second problem: all seven existing tutorial links point at `serve/tutorials/deployment-serve-llm/content/<name>/README.html`, which redirects to the Serve index. Ray publishes these under `_collections/` — per `doc/source/serve/examples.yml`, the link is `../_collections/serve/tutorials/deployment-serve-llm/<name>/README`. Repointed all eight. ## Testing `curl -sL` on each of the seven repointed links returns HTTP 200 at the requested path, against a soft 404 before. The Nemotron link 404s until Ray picks the template up in its next docs sync — the same lag every new tutorial has. Signed-off-by: svc-template-updater <282975472+svc-template-updater@users.noreply.github.com> Co-authored-by: svc-template-updater <282975472+svc-template-updater@users.noreply.github.com>
workspace
Readme
notebook
I have run the template in workspace to make sure it's runnable.