The idea is familiar from language models. Instead of training a model from scratch for one narrow task, you take a model pretrained on an enormous and diverse corpus and apply it to something it has never seen. Time series foundation models apply this approach to forecasting. Pretrained on millions of series, they learn the general grammar of sequential data: daily and weekly patterns, trends, and how a series reacts to outside factors like the weather.
That pretraining is not purely abstract. LOTSA, one of the pretraining archives, spans around 27 billion observations, and about 59 % of them come from the energy domain. Our grids were certainly not in there, but other series that behave a lot like them very probably were.
Using such a model without any training on the target data is called zero-shot, and that is what makes these models interesting in low-data scenarios.
We tested three of them: Chronos-2 (Amazon), TimesFM-2.5 (Google) and Moirai-2 (Salesforce). Each was run zero-shot and after parameter efficient fine-tuning on the grid’s own data. Some variants add Xreg, a technique we took from TimesFM’s reference implementation: a small linear regression on the weather and calendar covariates, fitted to whatever the foundation model gets wrong. TimesFM-2.5 needs it to use covariates at all, because the model itself only reads the demand history. Chronos-2 can take covariates directly, so for it Xreg is an alternative route. Against the foundation models stood four baselines: AutoARIMA as the production reference, XGBoost as a strong machine learning model, a ridge regression (a regularised linear regression), and naive persistence, which just repeats the last known input value as a flat line. A model that cannot beat a flat line has no business being deployed.