LightReasoner: The Amateur's Guide to Expert Thought

How a tiny model's failures teach a giant model to reason with 99% less data.

• View on GitHub • More from HKUDS

A giant mechanical titan looking into a small mirror held by a tiny clockwork bird, revealing a glowing gear missing from the bird's chest.
The 'Small-to-Large' paradigm uses a tiny model as a diagnostic mirror to find the exact moments where expert reasoning is required.

We propose LightReasoner, a novel framework that leverages the behavioral di vergence between a stronger expert model (LLM) and a weaker amateur model (SLM).

Jingyuan Wang, Author, The University of Hong Kong · LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

Key Takeaways

The High Cost of the Obvious

Traditional Supervised Fine-Tuning wastes massive amounts of compute on trivial tokens. When a model predicts the word 'The' in a math proof, it learns nothing new. Yet standard training pipelines treat every token equally, optimizing across the entire sequence. This brute force approach requires vast datasets and industrial server clusters to yield marginal improvements in reasoning capabilities.

LightReasoner delivers superior performance with remarkable token efficiency - achieving consistent improvements in zero-shot pass@1 accuracy while dramatically reducing computational overhead by 90% in total time, 80% in sampled problems, and 99% in tuned tokens compared to traditional SFT.

HKUDS/LightReasoner GitHub Repository, Project Repository · LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

Finding the Critical Moment

LightReasoner flips the traditional distillation hierarchy. Instead of a large model pouring knowledge into a small model, it uses the small model as a diagnostic tool. By comparing how an 'Expert' model and an 'Amateur' model respond to the same prompt, the system identifies the exact tokens where their logic diverges. When the small model is confused but the large model is confident, the system flags a critical reasoning moment.

The divergence filter highlights only the tokens where the amateur model fails, isolating the expert's unique reasoning.

Training on the Gaps

The core of this process lives in the sampling and fine-tuning scripts. The system calculates the KL Divergence between the two models' token distributions. Rather than training the model to match a single correct text output, it uses these divergent distributions as contrastive labels. This forces the model to align its internal probability distribution with the expert's thought process during those specific bottleneck moments.

The Frugal Frontier

This surgical approach stands in stark contrast to the reinforcement learning methods popularized by massive models. By focusing only on the 1 percent of tokens that actually matter for complex thought, LightReasoner enables high-performance math reasoning fine-tuning on consumer hardware.

FeatureTraditional SFTLightReasoner
Tokens Tuned100%< 1%
Data SourceGround Truth LabelsModel Divergence
Primary GoalGeneral KnowledgeReasoning Proficiency

From the HKUDS Lab

Developed by researchers at the University of Hong Kong, LightReasoner is part of a broader ecosystem of efficient AI tools. It proves that strategic data selection can outperform exhaustive training regimes.

Portrait of Jingyuan Wang