Swift-Net: The Speech Separator That Refuses to Cheat

A real-time audio-visual model that uses mouth motion, causal convolutions, and SRUs to pull one voice out of a crowd without peeking into the future.

8 min read • View on GitHub • More from JusperLee

A crowded room with multiple overlapping speakers, rendered as layered black-ink silhouettes. One central speaker is marked by a lit mouth and a clean ribbon of speech being pulled away from the noise, which explains how Swift-Net isolates a target voice in real time with visual guidance.
Swift-Net treats latency as a design constraint, not a deployment detail.
Key Takeaways

Most speech separators are judged by their accuracy offline. Swift-Net is judged by a harsher rule: can it separate one voice before the next syllable exists? That difference changes the whole stack, from padding to fusion to the recurrent block inside the network.

Why real-time speech separation is harder than it sounds

A separator can look impressive on a benchmark and still fail in the room. Live speech is messy, brief, and overlapping, which means the model has to decide with partial evidence and no second chance. Swift-Net is built around that constraint, not around the convenience of a fully observed clip.

That is why the visual stream matters. Lip motion arrives at the same moment as the audio frame, so it can stabilize the target voice without borrowing from the future. In this repo, real-time behavior is not a deployment trick. It is the design brief.

No future leakage, by construction

The model only sees what has already happened, while the mouth cue enters at the current time step.

A close-up mechanical conveyor carries audio frames toward a rigid gate that cuts off the unseen future. A small mouth-motion linkage enters from the side and nudges the current frame into alignment, which explains how Swift-Net enforces causality inside the model.
Causality is not a training hint here. It is a hard boundary in the pipeline.
if self.causal:
    x = x[:, :, :-self.causal_padding]

That slice is the heart of the claim. It trims off the tail that would otherwise leak future information, so the model cannot quietly cheat during inference. The code makes the promise explicit instead of trusting the training setup to behave.

How mouth motion fills in the missing context

The visual stream is not a second opinion. It is a disambiguation signal. In the Lightning module, the video branch turns the mouth frames into an embedding with mouth_emb = self.video_model(mouth.type_as(wav)), then passes that cue into the audio model at the exact moment it is needed.

Two separate strips, one jagged like a waveform and one shaped by mouth motion, feed into a brass bottleneck machine that outputs a clean speech mask. The image explains that Swift-Net fuses audio and video into one separation decision instead of treating video as decoration.
Video is a control signal in Swift-Net, not a sidecar feature.

That matters because overlap is where audio-only models wobble. A lip shape can anchor the target speaker before the acoustic signal fully resolves, which is exactly what a causal separator needs. Swift-Net is not using video for novelty. It is using video to survive uncertainty.

Inside the repo: a modular research stack

The repository looks like a research system that expects to be extended. The model, the visual backbone, and the trainer are separated cleanly enough that a researcher can swap pieces without rewriting the whole project. That is a strong signal that the code was built for reproducibility, not for a one-off demo.

The config files reinforce that impression. They expose different fusion modes and convolution paths, which is what you want when the goal is to test a causal claim under different settings. In other words, the repo is not frozen around one trick. It is organized to let the trick be measured.

Why SRU matters in a model called Swift-Net

SRUs are the speed play. They trade some of the overhead of heavier recurrent layers for a structure that can run faster, which matters when the system has to keep up with live conversation. In a model called Swift-Net, that is not branding. It is the point.

The naming is unusually honest. The model is not chasing the biggest possible offline score at any cost, because a separator that is late is often useless. Here, throughput and temporal discipline are part of the accuracy story.

Swift-Net versus the alternatives

ApproachModalityFuture audioLive useMain strengthMain trade-off
Audio-only separatorAudioNoYes, if streamingSimple and fastLoses visual grounding in overlap
Non-causal AVSSAudio + videoYesNo, not truly liveUses broader contextAdds latency and future leakage
Swift-NetAudio + videoNoYesKeeps separation causalGives up hindsight to stay real-time

That comparison is the whole editorial argument in miniature. Audio-only systems are efficient but blind to mouth motion. Non-causal audio-visual systems can be stronger offline, but they pay for it with latency. Swift-Net sits in the narrow middle, where the model must stay causal and still use the face to solve the problem.

What this repo is really for

Swift-Net reads like publication-grade research infrastructure. The preprocessing scripts, the config layer, the Lightning wrapper, and the visual backbone zoo all point to one goal: make causal audio-visual separation easy to reproduce, compare, and extend. If you care about live captioning, hearing aids, or teleconferencing, that is the useful signal. The codebase is built to test the real-time claim, not just to advertise it.