ZAPBench: Forecasting a Zebrafish Brain Like a Time Series

Google Research’s benchmark turns whole-brain calcium recordings, stimulus streams, and neuron position into a testbed for neural prediction.

10 min read • View on GitHub • More from google-research

A zebrafish head in profile with its brain rendered as dense, cell-like traces that extend forward into a forecast horizon. The image shows that the repository treats neural activity as a future state to predict, not just a static anatomical object to inspect.
The article’s central inversion: the brain becomes a forecast problem.
Key Takeaways

A Brain You Can Forecast

Most neuroscience repos start with analysis. ZAPBench starts with a wager: if you treat a zebrafish brain as a multivariate time series, can a model predict what each neuron does next? That shift matters because it turns the benchmark from a descriptive atlas into a test of temporal reasoning.

The Zebrafish Activity Prediction Benchmark (ZAPBench) measures progress on the problem of predicting cellular-resolution neural activity throughout an entire vertebrate brain.

google-research/zapbench README, Project documentation · zapbench README

Why Build a Benchmark Instead of Just a Model?

A single model can be clever and still be impossible to compare. ZAPBench exists to make progress accumulative, with a fixed task, shared data splits, and a common evaluation frame. The project comes out of Google Research with collaborators at HHMI Janelia and Harvard, but its deeper move is cultural: it takes a difficult neuroscience problem and gives it the same kind of benchmark discipline that made computer vision and forecasting fields move faster.

What ZAPBench Actually Predicts

The benchmark is simple to state and hard to solve. The inputs are past calcium traces from about 70,000 neurons, external stimuli such as visual patterns or water currents, and spatial context for each cell. The target is the future activity of those same neurons, so the model has to learn dynamics, not just correlations.

The benchmark is based on a novel dataset containing 4d light-sheet microscopy recordings of more than 70,000 neurons in a larval zebrafish brain, along with motion stabilized and voxel-level cell segmentations of these data that facilitate development of a variety of forecasting methods.

ZAPBench ICLR paper, Research abstract · ICLR paper abstract

The benchmark contract in one view: past traces and covariates go in, future neural activity comes out, and different forecasting models compete on the same task.

The Model Zoo Is the Point

ZAPBench is deliberately model-agnostic. It does not force a bespoke neuroscience architecture; it lets ordinary forecasting models make their case on biological data. That makes the benchmark more useful than a one-off model release because it tells you whether familiar ideas still work when the target is a brain.

This is the quiet insight: the repo is not asking forecasting models to become neuroscientists. It is asking whether their inductive biases are good enough to survive in a domain with far more structure than a commodity dataset.

How the Repo Streams a Whole Brain Without Falling Over

The engineering is built around not loading what you do not need. TensorStore keeps the giant arrays remote and chunked, lazy loading pulls only the requested slices, and helpers like restrict_specs_to_somas narrow training to cell bodies instead of the entire recording volume. Condition offsets handle the experimental windows, while the modeling stack uses JAX and Flax to keep the compute side fast enough to match the data side.

A cabinet of labeled data drawers feeds into a narrow pipeline that leads toward a training loop. The image explains why the repository uses lazy loading and chunked storage, because the whole brain only becomes manageable when the system touches small slices at a time.
Whole-brain scale only works when the pipeline opens small drawers, not the whole cabinet.

Why Zebrafish Sits in the Gap

EEG is broader but blurrier. Connectomics is detailed but static. Generic forecasting benchmarks are clean but biologically empty. Zebrafish sits in the middle: a vertebrate, whole-brain system you can still resolve at cellular scale, with enough structure to make prediction meaningful.

BenchmarkWhat it measuresStrengthBlind spot
EEG datasetsAggregate electrical activity across regionsHuman-scale coverage with low overheadCoarse spatial resolution and little cell-level detail
ConnectomicsStructural wiring and synaptic layoutReveals anatomy and long-term organizationStatic snapshot, not living dynamics
Generic forecasting benchmarksAbstract time series patternsClean model comparison and easy scalingNo biological meaning
ZAPBenchCellular-resolution neural activity plus stimuli and positionWhole-brain, dynamic, standardized forecast taskNarrower species scope and a heavier data pipeline

That middle ground is the point. If the data were simpler, the benchmark would be uninteresting. If it were more complex, the field would never agree on a shared task. ZAPBench is useful precisely because it is hard in a way that is still measurable.

What This Benchmark Is Really For

The long game is not one leaderboard. It is a shared language for systems neuroscience. Once a brain can be phrased as a forecast problem, models from time-series analysis, representation learning, and video prediction can enter the same conversation.

That is the open-source move here. ZAPBench makes a living vertebrate brain legible to standard ML tooling, and in doing so it turns a narrow dataset into a common reference point for a field that has needed one.