The Rollout Bottleneck: Inside THUDM/slime

How the framework behind GLM-4 bridges Megatron-LM and SGLang to scale reinforcement learning.

9 min read • View on GitHub • More from THUDM

A massive industrial hourglass where gears in the top half filter through a silicon wafer bottleneck into smooth sand.
Scaling reinforcement learning requires treating generation as a high-performance serving problem.
Portrait of Zilin Zhu

The main reason is that FSDP has not been part of our internal experimental/training workflow, and this is unlikely to change in the near term. As a result, it has not been where most of our engineering attention goes.

— Zilin Zhu, Maintainer
Key Takeaways

The Rollout Problem

The bitter lesson of scaling reinforcement learning for large language models is that training speed is bottlenecked by generation speed. During the rollout phase, the model must sample thousands of responses to evaluate. If the training framework treats this generation step as an afterthought, expensive GPU clusters sit idle waiting for tokens to appear.

Most traditional frameworks force a compromise. They either optimize for the gradient updates of the training loop or they optimize for the high-throughput serving of the generation loop. Doing both simultaneously at the 100-billion parameter scale breaks conventional architectures.

Slime, developed by Tsinghua University's THUDM and Z.ai, rejects the compromise. It treats reinforcement learning not as a single script, but as a distributed system of specialized microservices. It delegates training to Megatron-LM and delegates generation to SGLang. The resulting orchestration completely eliminates the rollout bottleneck.

Decoupling the Engine

The core philosophy of Slime is physical and logical decoupling. Training and rollout are distinct services kept in sync via a central Data Buffer. This architecture allows teams to allocate hardware asymmetrically. A cluster might dedicate eight nodes to the heavy matrix math of Megatron-LM gradient updates, while two distinct nodes run SGLang workers to churn through continuous batching and RadixAttention.

For smaller operations, Slime supports colocation. When running on the same nodes, the framework intelligently swaps model weights back and forth to CPU memory, ensuring that the active phase always has maximum GPU VRAM available.

A decoupled architecture flow showing three main blocks. On the left

To address the long-tail issue in agentic settings, we adopt a fully asynchronous RL training approach, ensuring that rare or long-tail data does not block the overall training process.

— Chengxing Xie, Main Contributor

The Metadata Passthrough

Routing requests to an inference engine introduces a critical problem for reinforcement learning. Standard inference routers strip out the internal metadata required for training. They return the generated text, but they discard the log probabilities and the specific routing decisions made by Mixture of Experts (MoE) layers.

Slime solves this with a custom FastAPI proxy called the Slime Router. It acts as a metadata-preserving passthrough.

A close-up of a pneumatic tube transport system where a clear capsule containing a delicate blueprint bypasses a heavy mechanical crusher via a dedicated side-pipe.
The Slime Router preserves critical MoE routing metadata that standard inference gateways typically discard.

When the SGLang worker generates a response, the Slime Router captures the exact MoE routing paths used during that specific rollout. It then forces the Megatron-LM trainer to use those identical paths during the gradient update. This technique, called Rollout Routing Replay, prevents the catastrophic training instability that occurs when expert selection shifts between generation and evaluation.

Online Draft Models

To maximize SGLang's throughput, Slime leans heavily on speculative decoding. A smaller draft model rapidly guesses the next tokens, and the massive policy model verifies them in parallel. However, in a reinforcement learning environment, the policy model is constantly changing. A static draft model quickly becomes obsolete and verification rates plummet.

Two sculptors working on the exact same marble statue simultaneously. One rapidly chips away large chunks, while the second follows behind with fine chisels.
Online Multi-Token Prediction allows the draft model to evolve simultaneously with the primary policy model.

Slime implements Online Multi-Token Prediction (MTP). The framework trains the draft layers simultaneously alongside the primary actor. As the main policy model improves, the draft model evolves with it. This ensures the speculative decoding engine maintains high acceptance rates throughout the entire training lifecycle.

The RL Scaling Landscape

Slime operates in a highly competitive space alongside frameworks like ByteDance's Verl and OpenRLHF. While those tools offer broad compatibility across different backend engines, Slime is ruthlessly specific. It is an industrial-grade tool built specifically for Megatron-LM.

Framework Core Engine Architecture Approach Inference Backend
THUDM/slime Megatron-LM SGLang-Native, Asynchronous Buffer SGLang
Verl Ray + Megatron/FSDP Hybrid Engine, Ray-centric vLLM / Lightllm
OpenRLHF Ray + DeepSpeed Scheduling-Centric, PPO/DPO optimized vLLM

This strict focus recently culminated in the controversial decision to completely remove support for the Fully Sharded Data Parallel (FSDP) backend. For the maintainers, the engineering overhead of supporting multiple training paradigms diluted their focus on maximum scale.

By dropping FSDP and betting entirely on Megatron-LM and SGLang, Slime abandons the generalist market. Instead, it provides the precise, highly optimized architecture required to train the next generation of reasoning models.


Sources: Repository architecture and codebase analysis from THUDM/slime. Quotes sourced from public GitHub issue discussions.