1. Model Introduction
NVIDIA Nemotron-3-Nano-4B-BF16 is a densenemotron_h hybrid model — interleaved Mamba and attention blocks
with squared-relu FFNs, no RoPE, and a 262 144 max position. miles wires it via
the megatron.bridge AutoBridge path, so there is no torch_dist conversion
step: the AutoBridge constructs the full Megatron provider from the HF
config.json at load time, including all Mamba-specific fields
(mamba_num_heads, mamba_state_dim, hybrid_override_pattern, etc.).
Key highlights:
- Hybrid architecture: Mamba + attention layers (
nemotron_hfamily). - Bridge-mode load:
--megatron-to-hf-mode bridge— no separate Megatron checkpoint. - No RoPE:
--position-embedding-type none. - Vocab: 131 072 tokens, padded to a multiple of 128.
2. Supported Variants
3. Environment Setup
3.1 Download model + datasets
--data-dir and --model-dir defaults; point them elsewhere with the flags.
3.2 No torch_dist conversion
AutoBridge loads the HF checkpoint directly. Both --hf-checkpoint and
--ref-load point at the HF directory, and --megatron-to-hf-mode bridge
turns on the bridge code path:
4. Launch
4.1 Quick start
scripts/run_nemotron_3_nano.py covers both Nano variants; --model-name picks the
dense 4 B here and the MoE 30B-A3B on the sibling page.
The recipe targets 1 node × 8 GPU (H100/H200), default cell TP=2 PP=2, and it is a
10-step smoke test (--num-rollout). Checkpoints are written under --output-dir
(default /root/shared_data).
5. Recipe Configuration
5.1 Parallelism
The recipe ships a starting cell ofTP=2 PP=2. Other verified cells (10-step RL
smoke tests, max train/rollout logprob diff): TP=2, TP=4, PP=2, CP=2, TP=2×PP=2.
Change the tp / pp / cp fields of the recipe to switch.
--sequence-parallel is not enabled in the dense smoke recipe; activation
checkpointing is also off. Dense Nemotron-3-Nano has no expert parallelism.
5.2 Algorithm
GRPO with low-variance KL:5.3 Rollout & SGLang
5.4 Optimizer
GPU Adam in the smoke recipe (no--optimizer-cpu-offload). Switch on CPU Adam if
memory pressure rises.
5.5 Notable quirks
Fromscripts/models/nemotron-3-nano-4b.py and scripts/run_nemotron_3_nano.py:
- No
--spec: the AutoBridge synthesizes the Megatron spec from HF config. --position-embedding-type none(no RoPE).--vocab-size 131072 --make-vocab-size-divisible-by 128.--attention-backend auto(the Mamba layers select their own kernel; flash-only is not safe here).- Bridge load is required for hybrid
nemotron_h: the AutoBridge wiresmamba_num_heads,mamba_state_dim,hybrid_override_pattern. PP additionally needs miles’ PP-unwrap shim (already on thefeat/nemotron-gemma4-rlbranch).

