Skip to main content

1. Model Introduction

Qwen3-Next is Alibaba’s next-generation Qwen architecture, swapping classical attention for a hybrid Gated DeltaNet + Full Attention design. Key highlights:
  • Hybrid Attention: combines Gated DeltaNet (linear attention) with Full Attention to handle context lengths up to 262 K tokens efficiently.
  • Highly Sparse MoE: 80 B total / 3 B active per token — drastically reduces FLOPs per token without sacrificing model capacity.
  • Multi-Token Prediction (MTP): built-in MTP layer enables EAGLE-style speculative rollout out of the box.
  • HuggingFace-wrapped Megatron backend: miles loads the Qwen/Qwen3-Next-80B-A3B HF module as a Megatron stage without re-implementing GDN from scratch.

2. Supported Variants

3. Environment Setup

3.1 Required env vars

--topology 4node needs MASTER_ADDR for the ray fan-out; --topology single-node needs nothing. Paths are Typer flags: --model-dir (default /root/models) for the staged checkpoint, --data-dir (default /root/datasets) for the datasets, --output-dir (default /root/shared_data) for what the run writes. On a multi-node run all three must be on a shared FS.

3.2 Download model + datasets

3.3 HF → Megatron torch_dist conversion

4. Launch

4.1 Quick start

4.2 Multi-node fan-out

--topology 4node performs ssh fan-out internally over /root/mpi_rack_hostfile — set MASTER_ADDR on the head node and the launcher reaches out to the workers. Pass --no-join-ray-workers when the ray cluster is already complete. --topology single-node never fans out.

5. Recipe Configuration

5.1 Parallelism

Both layouts are one launcher, scripts/run_qwen3_next_80b_a3b.py, selected by --topology: 4node colocates the rollout engines on the training GPUs; single-node dedicates the remaining 2 GPUs of the node to rollout, which also shrinks every batch dimension.

5.2 Algorithm

Both topologies use GSPO (--advantage-estimator gspo --eps-clip 4e-4); --use-kl-loss is not passed.

5.3 Rollout & SGLang

4node enables EAGLE speculative rollout:
single-node drops the EAGLE block and uses --rollout-num-gpus-per-engine 2 --rollout-num-gpus 2 --sglang-mem-fraction-static 0.8 --sglang-ep-size 1.

5.4 Optimizer

Both topologies enable CPU Adam:

5.5 Notable quirks

  • Gated DeltaNet (GDN) is loaded via the HuggingFace bridge; miles doesn’t re-implement GDN in Megatron native code.

6. Pairs Well With