One-RL-to-See-Them-All: Teaching Vision-Language Models to See via Reinforcement
How MiniMax-AI unified perception and reasoning into a single RL pipeline using Dynamic IoU rewards.
- The One-RL-to-See-Them-All framework unifies visual perception and logical reasoning into a single reinforcement learning pipeline.
- Dynamic IoU rewards overcome the limitations of binary pass-fail systems by providing a continuous gradient for spatial learning.
- A reference-free GRPO implementation significantly reduces VRAM overhead and allows for more aggressive exploration during training.
- Decoupling the reward logic into a standalone FastAPI server prevents CPU-heavy calculations from stalling GPU training loops.
The "Reasoning" era of AI has largely been a text-only affair. Models like DeepSeek-R1 have proven that Reinforcement Learning (RL) can teach an LLM to think through complex math and logic problems. But when it comes to the physical world—finding a specific object in a cluttered room, or understanding the spatial relationship between two items—VLMs have remained stubbornly reliant on Supervised Fine-Tuning (SFT).
One-RL-to-See-Them-All, developed by MiniMax-AI, changes this. It is the first framework to successfully apply RL to both visual reasoning and visual perception simultaneously. By treating coordinate bounding boxes as "tokens of thought," the system bridges the gap between abstract logic and physical geometry.
The Geometry of a Guess
The fundamental problem with applying standard RL to computer vision is the nature of the reward. In math, an answer is either right or wrong. But in object detection, if a model predicts a bounding box that is one pixel off from the ground truth, a binary reward system gives it a zero. The model receives no signal indicating it was close, making it nearly impossible to learn spatial coordinates via exploration.
MiniMax-AI solves this with a Dynamic IoU (Intersection over Union) reward mechanism. Instead of a pass/fail grade, the model receives a continuous reward based on how much its predicted box overlaps with the actual object. This "partial credit" creates the smooth gradient necessary for RL algorithms to optimize spatial understanding.
Architecture: The V-Triune Framework
To train a model that can both solve calculus problems and locate a coffee cup in an image, you need a unified data structure. The project introduces V-Triune (Visual Triple Unified Reinforcement Learning), which standardizes diverse tasks—OCR, math, puzzles, and object detection—into a single "Sample-Level" format.
We propose V-Triune (Visual Triple Unified Reinforcement Learning), a unified Reinforcement Learning (RL) system designed to advance Vision-Language Models (VLMs). It enables VLMs to jointly learn and master both visual reasoning and perception tasks within a single training pipeline. Our model, Orsta, trained with this approach, demonstrates how one RL framework can empower VLMs to "See Them All", delivering significant performance boosts across a diverse range of visual tasks.
Eliminating the Reference Model
Training 32B parameter models with RL is notoriously memory-intensive. Standard Group Relative Policy Optimization (GRPO) requires a "reference model"—a frozen copy of the original model used to calculate a KL-divergence penalty, preventing the RL policy from deviating too far from its initial state.
The MiniMax team engineered a workaround: they removed the reference model entirely. By relying on other regularization techniques, they freed up massive amounts of VRAM. This "Reference-Free GRPO" allows for more aggressive spatial exploration during training, enabling the model to test wildly different bounding box coordinates without being penalized for deviating from its SFT baseline.
| Feature | Standard GRPO | Orsta Optimized GRPO |
|---|---|---|
| Memory Overhead | High (requires frozen reference model) | Low (Reference-Free) |
| Exploration | Constrained by KL-divergence penalty | Unconstrained spatial exploration |
| VRAM Usage | Typically requires 2x model size | ~1x model size (plus optimizer states) |
The Reward Server: Decoupling Logic from Compute
Calculating the Dynamic IoU for thousands of images, or parsing complex LaTeX outputs to verify math, is CPU-heavy. If this logic sat inside the main training loop, the expensive GPUs would sit idle waiting for the CPU to finish grading the model's homework.
The solution is a decoupled architecture. The repository relies heavily on a separate reward_serving_fastapi.py service. This FastAPI-based "Reward Server" handles the heavy lifting of COCO dataset evaluation and rule-based verification asynchronously, ensuring the GPU rollout workers (powered by vLLM) never stall.
The implications of One-RL-to-See-Them-All extend beyond benchmark scores. By proving that a single RL pipeline can teach an AI to both reason and perceive, MiniMax-AI has provided an open-source blueprint for the next generation of autonomous agents—models that don't just read about the world, but actively look at it.