← All posts / Models

Google's TimesFM-3 Forecasts Multivariate Time Series in a Single Pass

Google Research's 330M-parameter time-series foundation model generates full multivariate forecast horizons — targets, covariates, and 9 quantiles — in one forward pass, topping Gift-Eval, FEV-Bench, and Time.

Google's TimesFM-3 Forecasts Multivariate Time Series in a Single Pass

While the AI industry’s attention fixates on ever-larger language models, Google Research quietly shipped an update to a far less glamorous — but commercially critical — category: time-series forecasting. On August 31, 2026, the team behind TimesFM published TimesFM-3, a zero-shot foundation model for multivariate forecasting that generates an entire forecast horizon, uncertainty bands included, in a single forward pass.

It is a dense little release. In a field where every vendor demo leans on chat, TimesFM-3 is the kind of model that ends up inside retail demand planning, treasury desks, observability dashboards, and energy grids — the decision-infrastructure layer of the economy that never makes headlines.

The problem: real forecasting is multivariate

When TimesFM debuted in 2024, it popularized the idea of a time-series foundation model: a decoder-only transformer pre-trained on massive corpora of real and synthetic time series, capable of zero-shot forecasting on data it has never seen. Since then, Google reports adoption across retail, finance, observability, manufacturing, healthcare, and the natural sciences.

But there was a persistent gap. Through TimesFM-2.5 (released September 2025), the family was strictly univariate — each forecast depended only on the history of a single series. The real world rarely cooperates with that constraint. As Google’s own example puts it: forecasting ice cream sales from past sales alone ignores foot traffic, sales of cones and syrups, weather forecasts, planned promotions, and holidays. Most genuinely valuable forecasting problems involve multiple interrelated series plus auxiliary features that jointly shape the future.

TimesFM-3 closes that gap. It is natively pre-trained for multivariate forecasting and supports, out of the box:

  • Multiple targets — jointly forecast related series (say, different brands of ice cream), with both point and quantile forecasts for every target
  • Past covariates — features known only historically, like past foot traffic
  • Past-future (dynamic) covariates — known future events, such as scheduled promotions or weather forecasts, used to steer the prediction

Under the hood: alternating attention and one-shot decoding

The architecture stays faithful to the family’s decoder-only transformer DNA, with two innovations that matter.

Multivariate token construction. Time series are grouped into patches of 32 time steps, normalized per series to handle wildly different scales. For past-future covariates, TimesFM-3 applies a “lookahead” trick: each token concatenates the current patch with future patches, letting the model ingest upcoming known signals directly.

Alternating attention. The transformer stack operates as a 2D grid over (time × series). Causal temporal attention runs horizontally across time — strictly causal within each series to prevent leakage. Full variate attention runs vertically across series, letting any token see all other series at the same time step — this is where cross-series correlations (a promotion lifting one product’s sales at the expense of another) get learned. The two mechanisms alternate across the stack’s 20 layers (model dim 1280, 16 heads, per the Hugging Face model card).

Non-autoregressive decoding is the headline change. Previous TimesFM versions generated forecasts one patch at a time — autoregressively — which meant latency, compounding error, and compute cost that scaled with horizon length. TimesFM-3 adopts Contiguous Patch Masking: masked placeholder tokens are appended for the entire future horizon alongside observed context, and the model fills in every masked patch simultaneously in one forward pass. Past-future covariates stay visible in the horizon, feeding the model known future events. The payoff per forecast step: 9 quantiles from the 10th to the 90th percentile, a full probabilistic uncertainty profile without iterative loops.

Google’s worked example shows the difference concretely. Given a planned promotion schedule as a past-future covariate, TimesFM-3 learns the promotion→sales-lift relationship from history and anticipates roughly a 20% bump on each promotion day. A univariate baseline projects the weekly pattern forward, blind to the calendar.

Benchmarks: top rank across all three suites

Google evaluated TimesFM-3 on three comprehensive public benchmarks — Gift-Eval, FEV-Bench, and Time — against recent foundation models including the multivariate-capable Chronos-2 and Toto 2.0 families, plus its predecessor TimesFM-2.5. The result: the top-ranked model on all three, in both point and probabilistic metrics, among all pre-trained foundation models.

Two details in the evaluation deserve attention. First, each benchmark includes TimesFM-3 twice. In univariate mode — no covariates, no cross-series information, each series treated independently — it already matches or beats every competing replicable foundation model. Switch to full multivariate mode and it takes another leap. Second, the multivariate gains are not bought with scale: the model is just 330 million parameters, pre-trained on more than 1 trillion real and synthetic time points (GiftEvalPretrain minus FEV-Bench overlaps, Wikipedia pageviews through November 2023, Google Trends top queries through end of 2022, plus synthetic and augmented data). Efficiency was a design constraint, not an afterthought — and the single-pass decode is what makes 330M parameters deliver full-horizon, 9-quantile forecasts quickly.

The catch: a non-commercial license

Developers reaching for the weights should read the fine print. TimesFM 1.0 and 2.x shipped under Apache-2.0; TimesFM-3 is released under the TimesFM Non-Commercial License v1.0 — weights live on Hugging Face as google/timesfm-3.0-pytorch, with code on GitHub. For enterprises hoping to embed it in revenue-facing products, that is a materially different proposition than its predecessors, and it reopens space for Apache-licensed rivals (Amazon’s Chronos line among them) in commercial deployments. For research, evaluation, and internal experimentation, it is fully available.

A BigQuery integration is “landing in the coming weeks”; until then Google points users at TimesFM-2.5 via the AI.FORECAST command to get familiar with the workflow — with the implied promise that the commercial path will run through Google Cloud rather than raw weights.

Why it matters

The time-series foundation model category is quietly consolidating around a small set of players — Google (TimesFM), Amazon (Chronos), and a cluster of startups — and the frontier is moving from “decent zero-shot univariate forecasts” to “multivariate, covariate-aware, probabilistic, single-pass.” TimesFM-3 checks every box at once, at a parameter count that fits on modest hardware.

The single-pass decode may be the most underrated part. Forecasting is an operational workload — run constantly, at scale, across thousands of series. Cutting horizon generation from an iterative loop to one forward pass changes the cost curve of the whole category. Combined with quantile output for uncertainty-aware planning, TimesFM-3 is less a research demo than a blueprint for the forecasting layer of agentic enterprise systems: an agent that plans inventory, staffing, or spend needs exactly this kind of fast, calibrated, context-aware prediction.

TimesFM-3 is available now on GitHub and Hugging Face. If your organization lives and dies by forecasts — and most do — it is worth a bench test before the BigQuery integration makes it frictionless.