Inside `google-research/flood-forecasting`: How Google Built a Flood Model for Broken Data

A deep dive into the code behind FloodHub, where basin embeddings, masked feature aggregation, and probabilistic outputs keep river forecasts useful even when the inputs are incomplete.

11 min read • View on GitHub • More from google-research

A bird's-eye editorial illustration of a watershed turned into a machine. Several sensor lines are broken or missing, but the forecasting engine still runs in the center. It explains the article's core idea: the model is built to keep working when real-world data feeds drop out.
Flood forecasting is not a clean-data problem. The code is designed to keep producing useful output when the river network, weather feeds, and basin records are incomplete.
Key Takeaways

The model that keeps talking when sensors go silent

The useful surprise in google-research/flood-forecasting is not that it predicts floods. It is that the model is built to survive the ordinary failures of flood data: missing gauges, delayed weather products, and basins with almost no local history.

That matters because flood forecasting is rarely a tidy forecasting benchmark. The hardest basins are often the least instrumented, which means the system has to reason from partial evidence and still produce something operationally useful.

From NeuralHydrology fork to FloodHub engine

The repository is a fork of NeuralHydrology, but the fork is the story. Google did not just add a new model file, then call it done. It reshaped the library around operational forecast sequences, data connectors, and training logic that fit FloodHub rather than a one-off research notebook.

With these AI-based technologies we extended the reliability of currently-available global nowcasts, on average, from zero to five days, and improved forecasts across regions in Africa and Asia to be similar to what are currently available in Europe.

Yossi Matias and Grey Nearing, VP Engineering & Research, and Research Scientist, Google Research · Google Research blog

Why basin DNA matters more than perfect weather data

The model's quiet superpower is its static context. Soil, slope, topography, and land cover get embedded into a latent representation, so the network can infer how a basin behaves even when the dynamic inputs are thin.

That is the real answer to the ungauged watershed problem. The model is not pretending to know every river directly. It is learning the signature of a basin, then using that signature to make a better guess when the live feeds are incomplete.

A close-up cross-section of a watershed with layered rock, soil, and contour rings folding into a compact latent core. It explains how static basin attributes become a learned representation that stands in for local history when gauges are sparse.
Static attributes do not replace measurements. They give the model a physical prior that makes sparse basins legible.

The masked-embedding trick, step by step

The newer model, MeanEmbeddingForecastLSTM, is the sharpest expression of the repo's main idea. Each source family gets its own embedding path, the static basin context is added separately, and a masked mean combines only the inputs that actually exist.

The model does not require every source to be present. It embeds what it has, masks what it lacks, and keeps the forecast path alive.

sources = [era5, chirps, gauges, satellite]
embeddings = []

for source in sources:
    if source.is_present:
        embeddings.append(source_embedding(source))

x = masked_mean(embeddings)
x = concat(x, static_basin_embedding)
h = lstm_stack(x)
y = output_head(h)

Why the newer model is more forgiving than the older one

The earlier HandoffForecastLSTM is a cleaner structural idea. A hindcast LSTM processes the past, a handoff network transforms its state, and a forecast LSTM takes over at the issue time. It is elegant, but it still assumes the input stream arrives in a way the system can hand off cleanly.

ModelCore trickWhen data is messyWhat it optimizes for
HandoffForecastLSTMTransfers state from hindcast to forecast through a learned handoff networkStrong when the sequence boundary is clean and continuousA precise transition between past and future
MeanEmbeddingForecastLSTMEmbeds each source family and masked-averages only the sources that existKeeps working when some feeds vanishOperational resilience under missing inputs
NeuralHydrology upstreamGeneral hydrological deep learning frameworkNot tuned for FloodHub's production forecast sequenceA research library that the fork specializes

Why uncertainty is part of the product

The repo's head.py makes the endpoint just as important as the encoder. A regression head gives a point estimate, but the CMAL head, short for Countable Mixture of Asymmetric Laplacians, turns the forecast into a distribution with tail behavior that better matches floods.

That is the right shape for this problem. Floods are not just about average error. They are about the consequences of being wrong in the tail, where a narrow river becomes a threat in hours, not days.

What this repo really teaches

The lesson is larger than hydrology. In physical systems, missingness is not an edge case. It is the default, and the best production models are the ones that can degrade gracefully without turning every gap into a failure.