SonicSim: The Open-Source Simulator That Teaches AI How Moving Sound Actually Behaves
A synthetic audio stack that treats motion as the signal, not metadata, then turns that motion into a dataset and benchmark for speech enhancement and separation.
- SonicSim's sharpest idea is to blend neighboring room impulse responses so a moving source sounds continuous instead of teleported.
- The repo is valuable because it links scanned scenes, material-aware acoustics, dry audio assets, and dataset generation in one pipeline.
- The benchmark stack matters because it turns synthetic audio into a repeatable test bed for enhancement models.
- SonicSim occupies a narrow but important niche between general embodied simulators and simpler room acoustics libraries.
Most synthetic audio tools freeze the world. SonicSim does the opposite. Its moving-source pipeline keeps acoustics continuous as a speaker walks, turns, or drifts through a room, which is the difference between a usable dataset and a neat demo.
That distinction matters because real listening is never static. A mic moves. A source pivots. Reflections change with every meter of travel, and models trained on frozen scenes learn the wrong lesson.
SonicSim interpolates the room, not just the waveform
The technical center of the repo lives in SonicSim_moving.py. Instead of treating motion like a sequence of disconnected room impulse responses, it interpolates between neighboring responses as the source advances, then keeps the output audible as one continuous trajectory.
That is a good engineering trade. SonicSim is not trying to re-render physics at audio-sample granularity. It is trying to preserve continuity, avoid clicks, and keep the motion believable enough for enhancement and separation models to learn from it.
We introduce SonicSim, a synthetic toolkit designed to generate highly customizable data for moving sound sources. SonicSim is developed based on the embodied AI simulation platform, Habitat-sim, supporting multi-level parameter adjustments, including scene-level, microphone-level, and source-level, thereby generating more diverse synthetic data.
That three-level control matters. If a dataset only changes the room, it misses microphone placement. If it only changes the microphone, it misses trajectories. SonicSim covers scene-level, microphone-level, and source-level variation in one stack.
Why the room matters as much as the voice
SonicSim_rir.py is the acoustic foundation. It maps semantic scene labels to material properties through mp3d_material_config.json, then uses those properties to shape reflections. That is why a rug, a couch, and a glass wall do not sound interchangeable.
The repo also stays flexible on output. It supports mono, binaural, and ambisonics, which makes it useful beyond one lab setup. The point is not just realism, but reuse.
The repo is a simulator, a dataset, and a benchmark
SonicSim is really two systems in one. The generator side can build SonicSet from sources like LibriSpeech, FSD50K, and FMA across many Matterport3D scenes. The benchmark side, especially enhancement/look2hear/, makes the result scientifically useful instead of merely plentiful.
That benchmark layer is the quiet strength of the repo. A generator can hide its own biases. A benchmark exposes them, because every model has to face the same task, the same metrics, and the same synthetic conditions.
The project is also built for speech people, not just simulation people. That is why the evaluation stack includes enhancement models and quality metrics such as DNSMOS and SIGMOS. It turns a room simulator into an experimental environment.
Where SonicSim fits among acoustic tools
| Project | Best for | Dynamic source handling | Scene realism | Speech benchmark layer |
|---|---|---|---|---|
| SonicSim | Dynamic speech data and moving-source experiments | Native interpolation pipeline | High, with Matterport3D and material mapping | Yes, through enhancement/look2hear |
| SoundSpaces | Audio-visual embodied AI and navigation | Native, but geared to navigation tasks | High, Habitat-based | Not the focus |
| Pyroomacoustics | Fast room acoustics prototyping | Mostly scripted and simpler | Moderate, geometric rooms | No built-in benchmark stack |
| NVIDIA Isaac Sim | Robotics and digital twins | Native, broad physics engine | Very high, but general-purpose | No dedicated speech benchmark |
The comparison is not about winners. It is about fit. SonicSim is narrower than Isaac Sim, but it is much more speech-specific. It is more scene-aware than a simple room acoustics library, and more benchmark-ready than a general embodied simulator.
That niche is exactly why it matters. The hard problem in moving-sound research is not just making audio that sounds plausible. It is making audio that is plausible, controlled, and repeatable enough to train models on and compare fairly.
SonicSim's real contribution is simple to say and hard to build: it treats motion as a first-class acoustic signal. Once you do that, the research question changes from whether a model can separate speech in a room to whether it can handle speech as a moving event in a real place.