Skip to main content
Miles ships RL recipes for every GLM generation currently in production: the GLM4.5 MoE at 106 B-A12B and 355 B-A32B, the compact GLM4.7 Flash with 64 routed experts, the 744 B-A40B GLM5 and GLM5.2 flagships, and GLM-5.3-Flash, a KDA + DSA hybrid on a different architecture entirely.

Variants

Fastest path to train

GLM4.7 Flash on a single 8× H100 node — the smallest GLM recipe:
See the GLM4.7 Flash page for weight conversion and the full walkthrough.

Which variant do I pick?

  • Single-node GLM first try → GLM4.7 Flash (glm4-7-flash).
  • MoE on a budget → GLM4.5-106B-A12B (glm4-5).
  • Full MoE scale (multi-node) → GLM4.5-355B-A32B (glm4-5).
  • Compact MoE for routing experiments (R3) → GLM4.7 Flash (glm4-7-flash).
  • Frontier scale (744 B) → GLM5.2 (glm5-2); GLM5/GLM5.1 (glm5) for the previous generation.
  • Hybrid linear + sparse attention with mHC → GLM-5.3-Flash (glm5-3-flash).