SonicSim: The Open-Source Simulator That Teaches AI How Moving Sound Actually Behaves

A synthetic audio stack that treats motion as the signal, not metadata, then turns that motion into a dataset and benchmark for speech enhancement and separation.

9 min read · JusperLee/SonicSim

A furnished room with a source moving along a path while reflected sound waves shift shape across the walls, floor, and ceiling. The image explains that SonicSim models acoustics as something continuous, so the sound changes as the source moves instead of jumping from one static snapshot to another.
SonicSim's premise is simple, but the implementation is not. Let motion reshape the room's sound field continuously.
Key Takeaways

Most synthetic audio tools freeze the world. SonicSim does the opposite. Its moving-source pipeline keeps acoustics continuous as a speaker walks, turns, or drifts through a room, which is the difference between a usable dataset and a neat demo.

That distinction matters because real listening is never static. A mic moves. A source pivots. Reflections change with every meter of travel, and models trained on frozen scenes learn the wrong lesson.

SonicSim interpolates the room, not just the waveform

The technical center of the repo lives in SonicSim_moving.py. Instead of treating motion like a sequence of disconnected room impulse responses, it interpolates between neighboring responses as the source advances, then keeps the output audible as one continuous trajectory.

Scrub the path to see how SonicSim blends neighboring room impulse responses instead of letting the source jump between positions.

That is a good engineering trade. SonicSim is not trying to re-render physics at audio-sample granularity. It is trying to preserve continuity, avoid clicks, and keep the motion believable enough for enhancement and separation models to learn from it.

We introduce SonicSim, a synthetic toolkit designed to generate highly customizable data for moving sound sources. SonicSim is developed based on the embodied AI simulation platform, Habitat-sim, supporting multi-level parameter adjustments, including scene-level, microphone-level, and source-level, thereby generating more diverse synthetic data.

Kai Li, et al., Authors · SonicSim

That three-level control matters. If a dataset only changes the room, it misses microphone placement. If it only changes the microphone, it misses trajectories. SonicSim covers scene-level, microphone-level, and source-level variation in one stack.

Why the room matters as much as the voice

A close view of one wall section in a scanned room, with different materials shown as tactile surfaces that change how sound reflects. The image explains that SonicSim depends on material-aware acoustics, not just room geometry, so a rug, a sofa, glass, and bare wall do not behave the same way.
SonicSim is not a shoebox simulator. The room has texture, and texture changes the sound.

SonicSim_rir.py is the acoustic foundation. It maps semantic scene labels to material properties through mp3d_material_config.json, then uses those properties to shape reflections. That is why a rug, a couch, and a glass wall do not sound interchangeable.

The repo also stays flexible on output. It supports mono, binaural, and ambisonics, which makes it useful beyond one lab setup. The point is not just realism, but reuse.

The repo is a simulator, a dataset, and a benchmark

Three stacked modules inside a machine-like frame: simulation at the top, synthetic dataset generation in the middle, and benchmarking at the bottom. The image explains that SonicSim is not just a generator, but a research stack that carries a signal from scene to dataset to model evaluation.
SonicSim works like a small research lab in a box. Generate, train, test, repeat.

SonicSim is really two systems in one. The generator side can build SonicSet from sources like LibriSpeech, FSD50K, and FMA across many Matterport3D scenes. The benchmark side, especially enhancement/look2hear/, makes the result scientifically useful instead of merely plentiful.

That benchmark layer is the quiet strength of the repo. A generator can hide its own biases. A benchmark exposes them, because every model has to face the same task, the same metrics, and the same synthetic conditions.

The project is also built for speech people, not just simulation people. That is why the evaluation stack includes enhancement models and quality metrics such as DNSMOS and SIGMOS. It turns a room simulator into an experimental environment.

Where SonicSim fits among acoustic tools

ProjectBest forDynamic source handlingScene realismSpeech benchmark layer
SonicSimDynamic speech data and moving-source experimentsNative interpolation pipelineHigh, with Matterport3D and material mappingYes, through enhancement/look2hear
SoundSpacesAudio-visual embodied AI and navigationNative, but geared to navigation tasksHigh, Habitat-basedNot the focus
PyroomacousticsFast room acoustics prototypingMostly scripted and simplerModerate, geometric roomsNo built-in benchmark stack
NVIDIA Isaac SimRobotics and digital twinsNative, broad physics engineVery high, but general-purposeNo dedicated speech benchmark

The comparison is not about winners. It is about fit. SonicSim is narrower than Isaac Sim, but it is much more speech-specific. It is more scene-aware than a simple room acoustics library, and more benchmark-ready than a general embodied simulator.

That niche is exactly why it matters. The hard problem in moving-sound research is not just making audio that sounds plausible. It is making audio that is plausible, controlled, and repeatable enough to train models on and compare fairly.

SonicSim's real contribution is simple to say and hard to build: it treats motion as a first-class acoustic signal. Once you do that, the research question changes from whether a model can separate speech in a room to whether it can handle speech as a moving event in a real place.