tiny-router: The 60KB Traffic Controller for LLM Agents

Stop wasting billions of parameters on basic intent. How a multi-head classifier uses interaction history and calibrated confidence to gatekeep expensive model calls.

7 min read • View on GitHub • More from UdaraJay

A massive golden sledgehammer swinging toward a tiny walnut, intercepted by a small mechanical hand.
Routing simple intents with a 70B parameter model is an expensive overkill. Specialized classifiers act as the gatekeeper to the computation layer.
Key Takeaways

The Sledgehammer and the Nut

Modern AI agents suffer from an acute sledgehammer problem. When a user replies "Yes" to a prompt, the system often spins up a frontier model, burns hundreds of milliseconds, and consumes expensive tokens just to decide what state to enter next. This is the LLM tax. High-performance systems cannot afford to route traffic using monolithic reasoning engines.

The open-source repository UdaraJay/tiny-router solves this by moving state management out of the prompt and into a dedicated routing layer. It is a specialized, multi-head text classifier designed to sit in front of heavier LLMs. Weighing in at roughly 60KB of Python and relying heavily on optimized Hugging Face encoders, it acts as an intelligent traffic controller. It categorizes short user messages into actionable dimensions, allowing the system to handle trivial interactions locally.

The core philosophy is simple. You do not need a generalized reasoning engine to know if a message is urgent. You need a fast, deterministic, and highly calibrated classifier.

Four Heads are Better Than One

Standard classification models return a single label. A user says "Book it," and the model categorizes it as an "action." But real-world routing requires more nuance. To effectively gatekeep an LLM, a system must understand multiple dimensions of an interaction simultaneously.

The architecture of tiny-router utilizes multi-task learning (MTL). Instead of a single output vector, the neural network terminates in four distinct classification heads: relation to the previous message, actionability, retention requirements, and urgency. These heads are not isolated. The architecture features dependency-aware routing where predictions from earlier heads feed directly into later ones. If the model determines a message is highly actionable, that probability shifts the baseline expectation for urgency.

Interactive multi-head decision tree. A central 'User Message' node splits into four distinct paths: Relation

By modeling the joint distribution of these traits, the router avoids the combinatorial explosion of creating separate labels for every possible state. It treats the interaction as a matrix of properties rather than a single bucket.

Beyond the Text: The Fusion Layer

A raw text string is rarely enough to determine intent. "Yes" means something entirely different after a confirmation prompt than it does after a cancellation warning. To solve this, the router relies on a Hybrid Input Fusion layer.

The model wraps a DeBERTa-v3-small encoder to handle the natural language. But before making a prediction, it concatenates the linguistic hidden state with structured metadata. This environment context is passed as a structured object containing the previous action, the previous outcome, and a critical recency_seconds metric.

A funnel blending a letter, a stopwatch, and a receipt into a single glowing rope.
Text alone is rarely enough. The fusion layer embeds recency and previous outcomes directly into the semantic representation.

This metadata is embedded and passed through a multi-layer perceptron before merging with the text vector. The result is a context-aware classification. The model inherently understands that repeating a command after a failed outcome implies a different routing priority than a fresh request.

The Honest Classifier

Neural networks are notoriously overconfident. A raw Softmax output might claim a 95% probability of a specific route, even when the model is entirely confused. In a production automation system, misplaced confidence leads to catastrophic misrouting.

The repository ships with a dedicated calibration pipeline to fix this. It utilizes temperature scaling to align the model's output probabilities with reality. A post-training script performs a bounded search to find the exact temperature that minimizes Expected Calibration Error (ECE) on a validation set.

Metric Zero-Shot LLM Routing tiny-router (ONNX)
Latency ~400ms - 800ms < 5ms
Cost (per 1M calls) ~$5.00+ Effectively $0 (Local compute)
Confidence Metric Hallucinated / Unreliable Calibrated Probability (ECE)
Context Awareness Requires heavy prompt injection Native metadata fusion

This calibration guarantees that when the model reports an 85% confidence score, it is correct exactly 85% of the time. This enables developers to set strict "Automation-Safe" thresholds. If the confidence clears the bar, the system routes the request instantly. If it falls short, it falls back to the heavy LLM.

A calibration slider demonstrating 'Automation-Safe' thresholds. A horizontal slider labeled 'Confidence Threshold' ranges from 0% to 100%. As the user moves the slider to the right (higher threshold)

From PyTorch to Production

Training a specialized model is only half the battle. Deploying it without dragging along gigabytes of PyTorch dependencies is the real challenge. The project treats deployment as a first-class citizen with an automated ONNX export pipeline.

Because ONNX handles dictionary outputs poorly, the export script wraps the model in a deterministic flattener. It maps the dynamic multi-head outputs into a strict tuple of tensors. The graph is then subjected to INT8 dynamic quantization.

A mechanical brain passing through narrowing glass rings to become a dense marble.
ONNX quantization maps heavy PyTorch training weights into a lightweight production binary without losing the structural logic.

The final artifact is a highly compressed, standalone binary. It can be executed on edge devices, inside serverless functions, or directly in a Rust backend using the ONNX Runtime. By pushing state routing to the edge, applications can reserve their LLM budgets for tasks that actually require reasoning.


Sources: