1. Model Introduction
Qwen3-Next is Alibaba’s next-generation Qwen architecture, swapping classical attention for a hybrid Gated DeltaNet + Full Attention design. Key highlights:- Hybrid Attention: combines Gated DeltaNet (linear attention) with Full Attention to handle context lengths up to 262 K tokens efficiently.
- Highly Sparse MoE: 80 B total / 3 B active per token — drastically reduces FLOPs per token without sacrificing model capacity.
- Multi-Token Prediction (MTP): built-in MTP layer enables EAGLE-style speculative rollout out of the box.
- HuggingFace-wrapped Megatron backend: miles loads the
Qwen/Qwen3-Next-80B-A3BHF module as a Megatron stage without re-implementing GDN from scratch.
2. Supported Variants
3. Environment Setup
3.1 Required env vars
--topology 4node needs MASTER_ADDR for the ray fan-out; --topology single-node needs nothing. Paths are Typer flags: --model-dir (default /root/models) for the staged checkpoint, --data-dir (default /root/datasets) for the datasets, --output-dir (default /root/shared_data) for what the run writes. On a multi-node run all three must be on a shared FS.
3.2 Download model + datasets
3.3 HF → Megatron torch_dist conversion
4. Launch
4.1 Quick start
4.2 Multi-node fan-out
--topology 4node performs ssh fan-out internally over /root/mpi_rack_hostfile — set MASTER_ADDR on the head node and the launcher reaches out to the workers. Pass --no-join-ray-workers when the ray cluster is already complete. --topology single-node never fans out.
5. Recipe Configuration
5.1 Parallelism
Both layouts are one launcher,scripts/run_qwen3_next_80b_a3b.py, selected by --topology:
4node colocates the rollout engines on the training GPUs; single-node dedicates the remaining 2 GPUs of the node to rollout, which also shrinks every batch dimension.
5.2 Algorithm
Both topologies use GSPO (--advantage-estimator gspo --eps-clip 4e-4); --use-kl-loss is not passed.
5.3 Rollout & SGLang
4node enables EAGLE speculative rollout:
single-node drops the EAGLE block and uses --rollout-num-gpus-per-engine 2 --rollout-num-gpus 2 --sglang-mem-fraction-static 0.8 --sglang-ep-size 1.
5.4 Optimizer
Both topologies enable CPU Adam:5.5 Notable quirks
- Gated DeltaNet (GDN) is loaded via the HuggingFace bridge; miles doesn’t re-implement GDN in Megatron native code.

