TIGER Makes Speech Separation Small Enough to Matter

A look at how this Tsinghua-built repo uses time-frequency interleaving and gain extraction to chase strong separation without the usual model bloat.

10 min read · JusperLee/TIGER

A tiny tiger-shaped mechanism stands between a tangled cloud of overlapping sound ribbons and two clean output streams. The image explains the repo's core promise: separate complex audio while keeping the machine compact enough to feel practical, not industrial.
TIGER's story is not brute force. It is a small separator trying to keep the useful structure and throw away the waste.
Key Takeaways

Speech separation sounds abstract until you try to use it in a noisy room. The hard part is not just splitting overlapping voices. It is keeping enough context to recover the signal while staying light enough to run outside a lab server.

Why TIGER starts with efficiency

Most modern separators buy quality with scale. TIGER takes the opposite bet: process the mixture in the time-frequency domain, extract gains instead of reconstructing everything from scratch, and keep the architecture lean enough that deployment does not become the next research project. That is why the repo reads like an answer to a product question as much as a paper question.

A close-up of a room where a voice bounces off a wall, a sofa, and a half-open door, leaving one direct path and several weaker reflections. The scene explains why EchoSet matters: speech separation has to survive room physics, not just overlap on a clean benchmark.
Clean mixtures are convenient. Real rooms are stubborn.

EchoSet is the clue

The dataset story matters. EchoSet is not a clean overlap benchmark dressed up as realism. Its focus on occlusions, materials, and reverberation says the authors care about the sound of rooms, not just the algebra of mixtures. That changes the job. A separator now has to survive reflections, not only interference.

TIGER turns one mixture into a structured time-frequency problem, then applies selective processing before reconstructing the sources.

How TIGER works

Think of the model as a grid problem. The waveform is turned into a time-frequency map, usually through an STFT-style transform that makes frequency bins and time slices explicit. Local convolutions pick up nearby structure, while selective attention handles longer dependencies without turning every frame into a full quadratic conversation.

That is where the repo's named pieces matter. `ConvNormAct` handles local feature extraction, `MSA` and `F^3A` split long-range work across time and frequency, and `InjectionMultiSum` helps preserve details as information moves through the stack. The result is not just speed. It is a way to keep the model small without flattening the signal into mush.

audionet = getattr(look2hear.models, config["audionet"]["audionet_name"])
model = audionet(**config["audionet"]["kwargs"])

The repo's base layer reinforces the same discipline. `pad_to_appropriate_length` keeps input shapes aligned with the model's stride pattern, `PyTorchModelHubMixin` makes loading straightforward, and the training script turns model selection into a YAML choice instead of a code rewrite. That is boring infrastructure, which is exactly what good research code should be.

The same pattern extends into `inference_dnr.py`, where the model is used for dialogue, music, and effects at 44.1 kHz. That widens the story from speech-only separation to a more general stem-splitting engine.

What the comparison really is

DimensionHeavy separatorsTIGER
Model sizeLarge backbones and broad attentionSub-1M-class footprint in the small variant
Core strategyPredict more with more capacityInterleave time and frequency, then extract gains
Data storyCleaner benchmark mixturesEchoSet's room physics and occlusions
Deployment postureServer-first or offline pipelinesEdge-friendly and real-time-minded

The interesting contrast is not TIGER versus one named baseline. It is TIGER versus the old assumption that better separation must mean a bigger model. The repo argues that careful operator design can win back a surprising amount of compute budget without giving up the separation task itself.