One-RL-to-See-Them-All: Teaching Vision-Language Models to See via Reinforcement

How MiniMax-AI unified perception and reasoning into a single RL pipeline using Dynamic IoU rewards.

• View on GitHub • More from MiniMax-AI

A mechanical eye looking through a viewfinder at a target, guided by concentric ripples rather than a binary pass/fail mechanism.
Traditional RL fails in vision because a near-miss gets the same zero-reward as a complete failure. Dynamic IoU provides the gradient needed to learn.

Key Takeaways

The "Reasoning" era of AI has largely been a text-only affair. Models like DeepSeek-R1 have proven that Reinforcement Learning (RL) can teach an LLM to think through complex math and logic problems. But when it comes to the physical world—finding a specific object in a cluttered room, or understanding the spatial relationship between two items—VLMs have remained stubbornly reliant on Supervised Fine-Tuning (SFT).

One-RL-to-See-Them-All, developed by MiniMax-AI, changes this. It is the first framework to successfully apply RL to both visual reasoning and visual perception simultaneously. By treating coordinate bounding boxes as "tokens of thought," the system bridges the gap between abstract logic and physical geometry.

The Geometry of a Guess

The fundamental problem with applying standard RL to computer vision is the nature of the reward. In math, an answer is either right or wrong. But in object detection, if a model predicts a bounding box that is one pixel off from the ground truth, a binary reward system gives it a zero. The model receives no signal indicating it was close, making it nearly impossible to learn spatial coordinates via exploration.

Dynamic IoU translates spatial overlap into a continuous reward gradient, giving the model 'partial credit' for near misses.

MiniMax-AI solves this with a Dynamic IoU (Intersection over Union) reward mechanism. Instead of a pass/fail grade, the model receives a continuous reward based on how much its predicted box overlaps with the actual object. This "partial credit" creates the smooth gradient necessary for RL algorithms to optimize spatial understanding.

Architecture: The V-Triune Framework

To train a model that can both solve calculus problems and locate a coffee cup in an image, you need a unified data structure. The project introduces V-Triune (Visual Triple Unified Reinforcement Learning), which standardizes diverse tasks—OCR, math, puzzles, and object detection—into a single "Sample-Level" format.

A vintage weaving loom weaving different threads labeled with tasks like 'Calculus' and 'Detection' into a single fabric.
The V-Triune framework unifies disparate visual and logical tasks into a single training pipeline.

We propose V-Triune (Visual Triple Unified Reinforcement Learning), a unified Reinforcement Learning (RL) system designed to advance Vision-Language Models (VLMs). It enables VLMs to jointly learn and master both visual reasoning and perception tasks within a single training pipeline. Our model, Orsta, trained with this approach, demonstrates how one RL framework can empower VLMs to "See Them All", delivering significant performance boosts across a diverse range of visual tasks.

MiniMax-AI, Project Maintainer · Repository: MiniMax-AI/One-RL-to-See-Them-All

Eliminating the Reference Model

Training 32B parameter models with RL is notoriously memory-intensive. Standard Group Relative Policy Optimization (GRPO) requires a "reference model"—a frozen copy of the original model used to calculate a KL-divergence penalty, preventing the RL policy from deviating too far from its initial state.

The MiniMax team engineered a workaround: they removed the reference model entirely. By relying on other regularization techniques, they freed up massive amounts of VRAM. This "Reference-Free GRPO" allows for more aggressive spatial exploration during training, enabling the model to test wildly different bounding box coordinates without being penalized for deviating from its SFT baseline.

FeatureStandard GRPOOrsta Optimized GRPO
Memory OverheadHigh (requires frozen reference model)Low (Reference-Free)
ExplorationConstrained by KL-divergence penaltyUnconstrained spatial exploration
VRAM UsageTypically requires 2x model size~1x model size (plus optimizer states)

The Reward Server: Decoupling Logic from Compute

Calculating the Dynamic IoU for thousands of images, or parsing complex LaTeX outputs to verify math, is CPU-heavy. If this logic sat inside the main training loop, the expensive GPUs would sit idle waiting for the CPU to finish grading the model's homework.

The solution is a decoupled architecture. The repository relies heavily on a separate reward_serving_fastapi.py service. This FastAPI-based "Reward Server" handles the heavy lifting of COCO dataset evaluation and rule-based verification asynchronously, ensuring the GPU rollout workers (powered by vLLM) never stall.

By decoupling the reward calculation into a separate FastAPI service, the system prevents CPU-bound verification tasks from blocking GPU training.

The implications of One-RL-to-See-Them-All extend beyond benchmark scores. By proving that a single RL pipeline can teach an AI to both reason and perceive, MiniMax-AI has provided an open-source blueprint for the next generation of autonomous agents—models that don't just read about the world, but actively look at it.