ZAPBench: Forecasting a Zebrafish Brain Like a Time Series
Google Research’s benchmark turns whole-brain calcium recordings, stimulus streams, and neuron position into a testbed for neural prediction.
- ZAPBench turns whole-brain zebrafish activity into a forecasting task that standard models can attack.
- The repo matters because it standardizes inputs, covariates, and evaluation, which makes biological prediction comparable across methods.
- TensorStore and lazy loading are not implementation details here, they are what makes whole-brain scale practical enough to benchmark.
- ZAPBench sits between coarse brain signals and static anatomy, which is why it can bridge neuroscience and time-series ML.
A Brain You Can Forecast
Most neuroscience repos start with analysis. ZAPBench starts with a wager: if you treat a zebrafish brain as a multivariate time series, can a model predict what each neuron does next? That shift matters because it turns the benchmark from a descriptive atlas into a test of temporal reasoning.
The Zebrafish Activity Prediction Benchmark (ZAPBench) measures progress on the problem of predicting cellular-resolution neural activity throughout an entire vertebrate brain.
Why Build a Benchmark Instead of Just a Model?
A single model can be clever and still be impossible to compare. ZAPBench exists to make progress accumulative, with a fixed task, shared data splits, and a common evaluation frame. The project comes out of Google Research with collaborators at HHMI Janelia and Harvard, but its deeper move is cultural: it takes a difficult neuroscience problem and gives it the same kind of benchmark discipline that made computer vision and forecasting fields move faster.
What ZAPBench Actually Predicts
The benchmark is simple to state and hard to solve. The inputs are past calcium traces from about 70,000 neurons, external stimuli such as visual patterns or water currents, and spatial context for each cell. The target is the future activity of those same neurons, so the model has to learn dynamics, not just correlations.
The benchmark is based on a novel dataset containing 4d light-sheet microscopy recordings of more than 70,000 neurons in a larval zebrafish brain, along with motion stabilized and voxel-level cell segmentations of these data that facilitate development of a variety of forecasting methods.
The Model Zoo Is the Point
ZAPBench is deliberately model-agnostic. It does not force a bespoke neuroscience architecture; it lets ordinary forecasting models make their case on biological data. That makes the benchmark more useful than a one-off model release because it tells you whether familiar ideas still work when the target is a brain.
- NLinear strips the problem down to a brutally simple baseline by normalizing away level and leaning on trend.
- TiDE treats static neuron position and dynamic stimulus streams as first-class inputs.
- TSMixer alternates between time mixing and feature mixing, which makes it a useful general-purpose test.
This is the quiet insight: the repo is not asking forecasting models to become neuroscientists. It is asking whether their inductive biases are good enough to survive in a domain with far more structure than a commodity dataset.
How the Repo Streams a Whole Brain Without Falling Over
The engineering is built around not loading what you do not need. TensorStore keeps the giant arrays remote and chunked, lazy loading pulls only the requested slices, and helpers like restrict_specs_to_somas narrow training to cell bodies instead of the entire recording volume. Condition offsets handle the experimental windows, while the modeling stack uses JAX and Flax to keep the compute side fast enough to match the data side.
Why Zebrafish Sits in the Gap
EEG is broader but blurrier. Connectomics is detailed but static. Generic forecasting benchmarks are clean but biologically empty. Zebrafish sits in the middle: a vertebrate, whole-brain system you can still resolve at cellular scale, with enough structure to make prediction meaningful.
| Benchmark | What it measures | Strength | Blind spot |
|---|---|---|---|
| EEG datasets | Aggregate electrical activity across regions | Human-scale coverage with low overhead | Coarse spatial resolution and little cell-level detail |
| Connectomics | Structural wiring and synaptic layout | Reveals anatomy and long-term organization | Static snapshot, not living dynamics |
| Generic forecasting benchmarks | Abstract time series patterns | Clean model comparison and easy scaling | No biological meaning |
| ZAPBench | Cellular-resolution neural activity plus stimuli and position | Whole-brain, dynamic, standardized forecast task | Narrower species scope and a heavier data pipeline |
That middle ground is the point. If the data were simpler, the benchmark would be uninteresting. If it were more complex, the field would never agree on a shared task. ZAPBench is useful precisely because it is hard in a way that is still measurable.
What This Benchmark Is Really For
The long game is not one leaderboard. It is a shared language for systems neuroscience. Once a brain can be phrased as a forecast problem, models from time-series analysis, representation learning, and video prediction can enter the same conversation.
That is the open-source move here. ZAPBench makes a living vertebrate brain legible to standard ML tooling, and in doing so it turns a narrow dataset into a common reference point for a field that has needed one.