TIGER Makes Speech Separation Small Enough to Matter
A look at how this Tsinghua-built repo uses time-frequency interleaving and gain extraction to chase strong separation without the usual model bloat.
- TIGER's central move is to treat speech separation as a time-frequency masking problem that can be solved without a heavyweight backbone.
- The project's efficiency claim matters because a small footprint is what makes separation believable for edge devices and real products.
- EchoSet shifts the benchmark from clean overlaps to room physics, which is where many separators actually break down.
- The repo is organized like a reproducible lab, with YAML configs, Lightning training, and hub-friendly loading that make the research easier to reuse.
Speech separation sounds abstract until you try to use it in a noisy room. The hard part is not just splitting overlapping voices. It is keeping enough context to recover the signal while staying light enough to run outside a lab server.
Why TIGER starts with efficiency
Most modern separators buy quality with scale. TIGER takes the opposite bet: process the mixture in the time-frequency domain, extract gains instead of reconstructing everything from scratch, and keep the architecture lean enough that deployment does not become the next research project. That is why the repo reads like an answer to a product question as much as a paper question.
EchoSet is the clue
The dataset story matters. EchoSet is not a clean overlap benchmark dressed up as realism. Its focus on occlusions, materials, and reverberation says the authors care about the sound of rooms, not just the algebra of mixtures. That changes the job. A separator now has to survive reflections, not only interference.
How TIGER works
Think of the model as a grid problem. The waveform is turned into a time-frequency map, usually through an STFT-style transform that makes frequency bins and time slices explicit. Local convolutions pick up nearby structure, while selective attention handles longer dependencies without turning every frame into a full quadratic conversation.
That is where the repo's named pieces matter. `ConvNormAct` handles local feature extraction, `MSA` and `F^3A` split long-range work across time and frequency, and `InjectionMultiSum` helps preserve details as information moves through the stack. The result is not just speed. It is a way to keep the model small without flattening the signal into mush.
audionet = getattr(look2hear.models, config["audionet"]["audionet_name"])
model = audionet(**config["audionet"]["kwargs"])
The repo's base layer reinforces the same discipline. `pad_to_appropriate_length` keeps input shapes aligned with the model's stride pattern, `PyTorchModelHubMixin` makes loading straightforward, and the training script turns model selection into a YAML choice instead of a code rewrite. That is boring infrastructure, which is exactly what good research code should be.
The same pattern extends into `inference_dnr.py`, where the model is used for dialogue, music, and effects at 44.1 kHz. That widens the story from speech-only separation to a more general stem-splitting engine.
What the comparison really is
| Dimension | Heavy separators | TIGER |
|---|---|---|
| Model size | Large backbones and broad attention | Sub-1M-class footprint in the small variant |
| Core strategy | Predict more with more capacity | Interleave time and frequency, then extract gains |
| Data story | Cleaner benchmark mixtures | EchoSet's room physics and occlusions |
| Deployment posture | Server-first or offline pipelines | Edge-friendly and real-time-minded |
The interesting contrast is not TIGER versus one named baseline. It is TIGER versus the old assumption that better separation must mean a bigger model. The repo argues that careful operator design can win back a surprising amount of compute budget without giving up the separation task itself.