The Rollout Bottleneck: Inside THUDM/slime
How the framework behind GLM-4 bridges Megatron-LM and SGLang to scale reinforcement learning.
The main reason is that FSDP has not been part of our internal experimental/training workflow, and this is unlikely to change in the near term. As a result, it has not been where most of our engineering attention goes.
- Slime eliminates reinforcement learning bottlenecks by decoupling training in Megatron-LM from high-throughput generation in SGLang.
- The Slime Router preserves critical Mixture of Experts metadata to ensure training stability during rollout routing replay.
- Online Multi-Token Prediction synchronizes draft model updates with the primary policy to maintain high speculative decoding efficiency.
- The framework prioritizes industrial-scale performance by focusing exclusively on Megatron-LM and removing support for FSDP.
The Rollout Problem
The bitter lesson of scaling reinforcement learning for large language models is that training speed is bottlenecked by generation speed. During the rollout phase, the model must sample thousands of responses to evaluate. If the training framework treats this generation step as an afterthought, expensive GPU clusters sit idle waiting for tokens to appear.
Most traditional frameworks force a compromise. They either optimize for the gradient updates of the training loop or they optimize for the high-throughput serving of the generation loop. Doing both simultaneously at the 100-billion parameter scale breaks conventional architectures.
Slime, developed by Tsinghua University's THUDM and Z.ai, rejects the compromise. It treats reinforcement learning not as a single script, but as a distributed system of specialized microservices. It delegates training to Megatron-LM and delegates generation to SGLang. The resulting orchestration completely eliminates the rollout bottleneck.
Decoupling the Engine
The core philosophy of Slime is physical and logical decoupling. Training and rollout are distinct services kept in sync via a central Data Buffer. This architecture allows teams to allocate hardware asymmetrically. A cluster might dedicate eight nodes to the heavy matrix math of Megatron-LM gradient updates, while two distinct nodes run SGLang workers to churn through continuous batching and RadixAttention.
For smaller operations, Slime supports colocation. When running on the same nodes, the framework intelligently swaps model weights back and forth to CPU memory, ensuring that the active phase always has maximum GPU VRAM available.
To address the long-tail issue in agentic settings, we adopt a fully asynchronous RL training approach, ensuring that rare or long-tail data does not block the overall training process.
The Metadata Passthrough
Routing requests to an inference engine introduces a critical problem for reinforcement learning. Standard inference routers strip out the internal metadata required for training. They return the generated text, but they discard the log probabilities and the specific routing decisions made by Mixture of Experts (MoE) layers.
Slime solves this with a custom FastAPI proxy called the Slime Router. It acts as a metadata-preserving passthrough.
When the SGLang worker generates a response, the Slime Router captures the exact MoE routing paths used during that specific rollout. It then forces the Megatron-LM trainer to use those identical paths during the gradient update. This technique, called Rollout Routing Replay, prevents the catastrophic training instability that occurs when expert selection shifts between generation and evaluation.
Online Draft Models
To maximize SGLang's throughput, Slime leans heavily on speculative decoding. A smaller draft model rapidly guesses the next tokens, and the massive policy model verifies them in parallel. However, in a reinforcement learning environment, the policy model is constantly changing. A static draft model quickly becomes obsolete and verification rates plummet.
Slime implements Online Multi-Token Prediction (MTP). The framework trains the draft layers simultaneously alongside the primary actor. As the main policy model improves, the draft model evolves with it. This ensures the speculative decoding engine maintains high acceptance rates throughout the entire training lifecycle.
The RL Scaling Landscape
Slime operates in a highly competitive space alongside frameworks like ByteDance's Verl and OpenRLHF. While those tools offer broad compatibility across different backend engines, Slime is ruthlessly specific. It is an industrial-grade tool built specifically for Megatron-LM.
| Framework | Core Engine | Architecture Approach | Inference Backend |
|---|---|---|---|
| THUDM/slime | Megatron-LM | SGLang-Native, Asynchronous Buffer | SGLang |
| Verl | Ray + Megatron/FSDP | Hybrid Engine, Ray-centric | vLLM / Lightllm |
| OpenRLHF | Ray + DeepSpeed | Scheduling-Centric, PPO/DPO optimized | vLLM |
This strict focus recently culminated in the controversial decision to completely remove support for the Fully Sharded Data Parallel (FSDP) backend. For the maintainers, the engineering overhead of supporting multiple training paradigms diluted their focus on maximum scale.
By dropping FSDP and betting entirely on Megatron-LM and SGLang, Slime abandons the generalist market. Instead, it provides the precise, highly optimized architecture required to train the next generation of reasoning models.