TimesFM: Google’s Forecasting Model That Thinks Like an LLM

A decoder-only time-series foundation model that replaces per-dataset training with zero-shot prediction, patch-based decoding, and probabilistic forecasts.

8 min read • View on GitHub • More from google-research

A long stream of jagged time-series patches enters a decoder machine, then splits into several forecast paths on the far side. A side channel strips and restores scale before and after the model, showing that normalization is part of the pipeline, not an afterthought.
TimesFM treats forecasting like generation. The model ingests patches, decodes future values, and restores the original scale at the end.
Key Takeaways

The interesting part of TimesFM is not that Google made another forecasting model. It is that the repo behaves like an LLM stack that learned to speak in time steps instead of words. The model is built to forecast unseen series with no task-specific training, and that makes it feel less like a specialized predictor and more like a reusable forecasting layer.

TimesFM is a forecasting model, pre-trained on a large time-series corpus of 100 billion real world time-points, that displays impressive zero-shot performance on a variety of public benchmarks from different domains and granularities.

Rajat Sen and Yichen Zhou, Google Research · A decoder-only foundation model for time-series forecasting

The weird idea: forecasting as decoding

TimesFM starts from a blunt premise. If language models can generate the next token from context, a forecasting model can generate the next point from history. That shift matters because the unit of work is no longer a dataset-specific regressor. It is a pre-trained decoder that can be dropped onto a new series and asked to continue the pattern.

That is the zero-shot claim in plain English. You do not retrain TimesFM for every new product line, sensor fleet, or demand curve. You feed it the past, let it infer the structure, and ask for a distribution over what comes next.

Why TimesFM 2.5 matters more than TimesFM 2.0

FeatureTimesFM 2.5Why it matters
Parameters200MSmaller than the earlier 500M line, but easier to deploy and iterate on.
Context length16KIt can look much farther back without resorting to task-specific retraining.
Forecast outputQuantiles to 1K horizonThe model returns uncertainty bands, not just a single line.
Frequency handlingAutomaticThe model no longer needs a manual frequency token in the common path.
Inference styleRebuilt APIThe repo feels more like a production interface than a paper artifact.

The story here is efficiency, not bloat. TimesFM 2.5 cuts parameter count while expanding context, which is the kind of trade-off you expect from a mature system, not a novelty project. It is trying to become easier to run, easier to integrate, and more structurally robust.

A tight assembly line groups 32-point input tiles into larger blocks, sends them through a rotary encoder wheel, then expands them into 128-point output blocks. Nine forecast lanes branch outward in parallel, with the outer lanes enclosing the median line to show quantile uncertainty.
Patches turn a long series into manageable chunks. Quantiles turn a point guess into a forecast band.

The model’s DNA lives in config, not just weights

from timesfm.configs import ForecastConfig

cfg = ForecastConfig(
    force_flip_invariance=True,
    infer_is_positive=True,
    horizon=1000,
)

# The configuration is not just plumbing.
# It encodes output behavior the model should respect at inference time.

That small config object is doing real work. `force_flip_invariance` pushes the model toward symmetry. `infer_is_positive` keeps it from inventing negative demand or sales where that would make no sense. In other words, the repo bakes common-sense constraints into the output path.

This is one reason TimesFM reads as production-minded. It is not only learning from data. It is also being told what kinds of answers are physically or operationally admissible.

How the model turns a signal into tokens

The pipeline is closer to token generation than to classical regression. Scale handling, patching, decoding, and uncertainty output are all separate stages.

The patching step is the crucial translation layer. By grouping contiguous time points into blocks, TimesFM reduces a long numeric sequence into something the decoder can process more efficiently. That is the same trick that made patch-based vision and time-series models practical, but here it is wired into an autoregressive forecasting stack.

A foundation model for time-series forecasting, in contrast, can provide decent out-of-the-box forecasts on unseen time-series data with no additional training, enabling users to focus on refining forecasts for the actual downstream task like retail demand planning.

Rajat Sen and Yichen Zhou, Google Research · A decoder-only foundation model for time-series forecasting

The engine room: RoPE, RMSNorm, KV cache

Under the hood, the repo rolls its own Transformer primitives rather than leaning on a generic black box. Rotary positional embeddings help the model keep track of long context. RMSNorm gives the stack a modern normalization choice. The decode cache splits prefill from incremental generation, which is exactly the kind of detail you expect when the model is meant to behave like a decoder, not a one-shot regressor.

That is the LLM analogy made concrete. The model first absorbs context, then steps forward one forecast chunk at a time while reusing cached state. The architecture is not pretending to be language modeling. It is borrowing the most useful parts of that playbook.

RevIN is the quiet trick that makes the whole thing usable

RevIN is easy to overlook because it sounds like housekeeping. It is not. In practice, time-series data arrives on wildly different scales, and a model that ignores that problem becomes brittle fast. TimesFM normalizes the input, forecasts in that stabilized space, and then denormalizes the result back into the original units.

That reversible flow is what makes the system feel sane. The model can work across sales, traffic, weather, and sensors without pretending those signals live in the same numeric universe.

What it beats, what it borrows, and what it refuses to be

Model familyWhat it is good atWhere TimesFM differs
ARIMA / ETSStrong classical baselinesTimesFM avoids per-series fitting and aims for zero-shot use.
DeepARFlexible supervised forecastingTimesFM is pre-trained once and reused across domains.
PatchTSTStrong long-horizon transformer forecastingTimesFM borrows the patch idea but packages it as a foundation model.
llmtime-style LLM forecastingZero-shot promptingTimesFM is much smaller and purpose-built for time series.

The cleanest way to place TimesFM is as a reusable forecasting layer. It borrows the patch idea from transformer forecasting, the decoding mindset from LLMs, and the uncertainty framing from modern probabilistic forecasting. What it refuses to be is a one-off model that needs babysitting every time the data changes.

The bridge from research repo to enterprise workflow

BigQuery support matters because it changes the shape of the product story. This is not just a paper implementation sitting in a lab repo. It has an official path into a warehouse workflow, which means the same core idea can be used where business data already lives.

That bridge is the real signal. TimesFM is trying to move forecasting out of bespoke model projects and into infrastructure. Once that happens, the hardest part is no longer training a separate model for each team. It is deciding where a foundation forecast is good enough, and where human judgment still needs to take the wheel.