Swift-Net: The Speech Separator That Refuses to Cheat
A real-time audio-visual model that uses mouth motion, causal convolutions, and SRUs to pull one voice out of a crowd without peeking into the future.
- Swift-Net makes real-time speech separation a structural property by cutting off future context inside the model.
- The mouth embedding is not decorative, because it anchors the target speaker when audio alone is ambiguous.
- SRUs and causal padding are there to keep the system fast enough for live use, not just accurate in offline tests.
- The repository reads like research infrastructure for reproducible causal AVSS experiments, not a polished product shell.
Most speech separators are judged by their accuracy offline. Swift-Net is judged by a harsher rule: can it separate one voice before the next syllable exists? That difference changes the whole stack, from padding to fusion to the recurrent block inside the network.
Why real-time speech separation is harder than it sounds
A separator can look impressive on a benchmark and still fail in the room. Live speech is messy, brief, and overlapping, which means the model has to decide with partial evidence and no second chance. Swift-Net is built around that constraint, not around the convenience of a fully observed clip.
That is why the visual stream matters. Lip motion arrives at the same moment as the audio frame, so it can stabilize the target voice without borrowing from the future. In this repo, real-time behavior is not a deployment trick. It is the design brief.
No future leakage, by construction
if self.causal:
x = x[:, :, :-self.causal_padding]
That slice is the heart of the claim. It trims off the tail that would otherwise leak future information, so the model cannot quietly cheat during inference. The code makes the promise explicit instead of trusting the training setup to behave.
How mouth motion fills in the missing context
The visual stream is not a second opinion. It is a disambiguation signal. In the Lightning module, the video branch turns the mouth frames into an embedding with mouth_emb = self.video_model(mouth.type_as(wav)), then passes that cue into the audio model at the exact moment it is needed.
That matters because overlap is where audio-only models wobble. A lip shape can anchor the target speaker before the acoustic signal fully resolves, which is exactly what a causal separator needs. Swift-Net is not using video for novelty. It is using video to survive uncertainty.
Inside the repo: a modular research stack
The repository looks like a research system that expects to be extended. The model, the visual backbone, and the trainer are separated cleanly enough that a researcher can swap pieces without rewriting the whole project. That is a strong signal that the code was built for reproducibility, not for a one-off demo.
- `look2hear/models/`: SwiftNet and causal variants of other separators.
- `look2hear/videomodels/`: visual backbones such as ResNet, ShuffleNet, and FRCNN.
- `look2hear/layers/`: reusable normalization, RNN, and attention blocks.
- `look2hear/system/`: PyTorch Lightning wrappers for training and validation.
- `configs/` and `DataPreProcess/`: experiment settings and dataset formatting scripts.
The config files reinforce that impression. They expose different fusion modes and convolution paths, which is what you want when the goal is to test a causal claim under different settings. In other words, the repo is not frozen around one trick. It is organized to let the trick be measured.
Why SRU matters in a model called Swift-Net
SRUs are the speed play. They trade some of the overhead of heavier recurrent layers for a structure that can run faster, which matters when the system has to keep up with live conversation. In a model called Swift-Net, that is not branding. It is the point.
The naming is unusually honest. The model is not chasing the biggest possible offline score at any cost, because a separator that is late is often useless. Here, throughput and temporal discipline are part of the accuracy story.
Swift-Net versus the alternatives
| Approach | Modality | Future audio | Live use | Main strength | Main trade-off |
|---|---|---|---|---|---|
| Audio-only separator | Audio | No | Yes, if streaming | Simple and fast | Loses visual grounding in overlap |
| Non-causal AVSS | Audio + video | Yes | No, not truly live | Uses broader context | Adds latency and future leakage |
| Swift-Net | Audio + video | No | Yes | Keeps separation causal | Gives up hindsight to stay real-time |
That comparison is the whole editorial argument in miniature. Audio-only systems are efficient but blind to mouth motion. Non-causal audio-visual systems can be stronger offline, but they pay for it with latency. Swift-Net sits in the narrow middle, where the model must stay causal and still use the face to solve the problem.
What this repo is really for
Swift-Net reads like publication-grade research infrastructure. The preprocessing scripts, the config layer, the Lightning wrapper, and the visual backbone zoo all point to one goal: make causal audio-visual separation easy to reproduce, compare, and extend. If you care about live captioning, hearing aids, or teleconferencing, that is the useful signal. The codebase is built to test the real-time claim, not just to advertise it.