Apollo-data-preprocess: The tiny preprocessing repo that teaches Apollo what music is worth keeping

Apollo-data-preprocess is not a model repo. It is the filter that strips silence, slices tracks into learnable windows, and turns raw stems into HDF5 training fuel for high-fidelity music restoration.

11 min read • View on GitHub • More from JusperLee

A wide editorial illustration showing a river of tangled audio tape entering a filtering machine and emerging as neat stacked reels and storage blocks. The scene explains that preprocessing turns messy music files into training-ready datasets by removing dead air and standardizing the output.
Apollo's preprocessing layer is a gatekeeper, not a showcase. It decides which parts of a track deserve to become training data.
Key Takeaways

Most preprocessing scripts are invisible by design. This one deserves a closer look because it decides what Apollo hears, what it ignores, and how expensive each training step becomes.

The repository sits behind Apollo, the music restoration project, and its job is simple to describe but hard to do well: take raw music datasets like MUSDB18-HQ and MoisesDB, remove unhelpful silence, cut the rest into consistent windows, and serialize everything into HDF5 files that train efficiently.

During data preprocessing, we drew inspiration from music separation techniques and implemented the following steps: Source Activity Detection (SAD)... Data Augmentation... Simulating Dynamic Bitrate Compression... Rescaling... Saving as HDF5

JusperLee (Kai Li), Author/Maintainer · JusperLee/Apollo

What the repo actually does

The codebase is tiny on purpose. Two scripts, musdb_preprocess.py and moisesdb_preprocess.py, handle two dataset layouts and share nearly the same core logic. That is not a polished product pattern. It is a research pattern: keep the path from dataset to training file obvious.

At the center is a music-aware activity detector. It does not try to recognize words or speakers. It measures audio power in windows, estimates a track-specific floor from the non-silent regions, then keeps only 6 second segments where more than half the frames clear that threshold. In plain terms, it refuses to waste training budget on dead air.

The repo splits into dataset-specific entry points, then converges on one reusable preprocessing path. That convergence is the whole point.

The differentiator is not cleaning. It is selective hearing.

Most data pipelines optimize for uniformity. Apollo optimizes for density. Its VAD logic uses a percentile of each track's own energy distribution, which means a quiet classical passage and a compressed rock chorus are both judged relative to their internal dynamics, not an arbitrary global decibel line.

That matters because restoration models learn from examples, and examples with no musical content are expensive noise. If you are training a model to restore compressed music, you want the model to spend its capacity on transients, harmonics, and layered stems, not empty bars.

A close-up editorial illustration of a single audio waveform being measured by a mechanical gauge. Dense peaks are kept while a quiet tail is shaved away, showing how a relative threshold preserves musically useful regions and discards low-energy segments.
Apollo's threshold is adaptive, not absolute. It tracks the shape of each song instead of assuming all music behaves the same way.

Why the format choice matters

The output is HDF5, and that choice is doing more work than it first appears. Instead of leaving training to crawl through loose audio files, the scripts package segments into indexed arrays that can be loaded efficiently and repeatedly.

Each track becomes its own file, which is a research-friendly compromise. It is easy to parallelize, easy to inspect, and easy to rerun when a preprocessing tweak changes the sample population.

def VAD(x, win=441, step=22050, sr=44100):
    # Split into 6 second windows with 50% overlap
    # Estimate a track-specific noise floor from non-silent frames
    # Keep a window only if enough frames exceed the threshold
    pass

# Raw audio -> VAD filtering -> numpy array -> HDF5 dataset

That pattern fits the rest of the repository. The code is readable, direct, and not especially abstracted. It makes fewer promises than a production framework, but it also leaves fewer mysteries behind.

How it compares

DimensionApollo-data-preprocessTypical general-purpose prep
Primary goalKeep only musically dense segments for restoration trainingNormalize or transform data for broad reuse
ThresholdingRelative, track-aware percentile based VADOften fixed thresholds or generic heuristics
Input shapeMulti-track music datasets with stemsSingle streams, text, images, or mixed modalities
OutputPer-track HDF5 files for fast training accessVariable formats, often less optimized for audio loops
Trade-offLess DRY, more explicitMore generalized, less domain-specific

The comparison is not really between good and bad engineering. It is between a generic data tool and a domain-shaped tool. Apollo's scripts are narrow because music restoration is narrow: the training data has to preserve timing, energy, and spectral detail at 44.1 kHz, or the model learns from an inferior world.

What this says about the project

This repo is a marker of research maturity. The authors did not stop at model architecture. They also published the part that makes the benchmark reproducible, which is often where the real advantage hides.

It also tells you what Apollo values. The project is willing to spend preprocessing complexity to preserve audio quality later. That is a strong signal that the goal is restoration, not convenience.

The license reinforces that framing. This is an academic artifact, built for research use and sharing, not a commercial pipeline product. The repo is small, but its role is large: it shapes the data lens through which the model learns what compressed music should become.