LightReasoner: The Amateur's Guide to Expert Thought
How a tiny model's failures teach a giant model to reason with 99% less data.

We propose LightReasoner, a novel framework that leverages the behavioral di vergence between a stronger expert model (LLM) and a weaker amateur model (SLM).
- LightReasoner identifies critical reasoning steps by comparing the divergence between expert and amateur model predictions.
- The framework achieves high-performance math reasoning by fine-tuning on less than one percent of total tokens.
- Contrastive labels force models to align their internal probability distributions during specific logical bottlenecks.
- This surgical training approach reduces computational overhead by 90% compared to traditional supervised fine-tuning.
The High Cost of the Obvious
Traditional Supervised Fine-Tuning wastes massive amounts of compute on trivial tokens. When a model predicts the word 'The' in a math proof, it learns nothing new. Yet standard training pipelines treat every token equally, optimizing across the entire sequence. This brute force approach requires vast datasets and industrial server clusters to yield marginal improvements in reasoning capabilities.
LightReasoner delivers superior performance with remarkable token efficiency - achieving consistent improvements in zero-shot pass@1 accuracy while dramatically reducing computational overhead by 90% in total time, 80% in sampled problems, and 99% in tuned tokens compared to traditional SFT.
Finding the Critical Moment
LightReasoner flips the traditional distillation hierarchy. Instead of a large model pouring knowledge into a small model, it uses the small model as a diagnostic tool. By comparing how an 'Expert' model and an 'Amateur' model respond to the same prompt, the system identifies the exact tokens where their logic diverges. When the small model is confused but the large model is confident, the system flags a critical reasoning moment.
Training on the Gaps
The core of this process lives in the sampling and fine-tuning scripts. The system calculates the KL Divergence between the two models' token distributions. Rather than training the model to match a single correct text output, it uses these divergent distributions as contrastive labels. This forces the model to align its internal probability distribution with the expert's thought process during those specific bottleneck moments.
The Frugal Frontier
This surgical approach stands in stark contrast to the reinforcement learning methods popularized by massive models. By focusing only on the 1 percent of tokens that actually matter for complex thought, LightReasoner enables high-performance math reasoning fine-tuning on consumer hardware.
| Feature | Traditional SFT | LightReasoner |
|---|---|---|
| Tokens Tuned | 100% | < 1% |
| Data Source | Ground Truth Labels | Model Divergence |
| Primary Goal | General Knowledge | Reasoning Proficiency |
From the HKUDS Lab
Developed by researchers at the University of Hong Kong, LightReasoner is part of a broader ecosystem of efficient AI tools. It proves that strategic data selection can outperform exhaustive training regimes.