JuaAI/ts-icl-pretraining-corpus
TS-ICL Pretraining Corpus (community reconstruction) A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.
TS-ICL Pretraining Corpus (community reconstruction)
A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and normalized into a single schema.
Curated by: Jua — reconstruction, normalization, synthetic generation, validation and packaging.
Not an official release. This is a derivative dataset built by Jua for research reproducibility; it is not affiliated with or endorsed by the TS-ICL authors. Each subset retains the license of its upstream source (see Licensing).
At a glance
- 39 datasets, 1,399,719 series, 8,690,897,054 observations
- Sources: LOTSA (21), Chronos (10), TempoPFN synthetic (8)
- Single consistent schema, sharded Parquet (zstd)
Repository layout
data/<config>/*.parquet # one folder per dataset (= HF config)
meta/datasets.yaml # build registry (sources, fingerprints, subsampling)
meta/reports/<config>.json # per-dataset build/QA report
meta/qa.json # aggregated QA
QUALITY_REPORT.md # human-readable QA tableSchema
Every row is one (possibly multivariate) time series:
How to load
from datasets import load_dataset
# one dataset (config)
ds = load_dataset("<repo_id>", "nn5_weekly", split="train")
# everything, streamed
ds = load_dataset("<repo_id>", "all", split="train", streaming=True)Datasets (provenance)
Reconstruction methodology
- LOTSA subsets: from
Salesforce/lotsa_data(Arrow); targets passed through. - Chronos subsets: from
autogluon/chronos_datasets; per-seriestimestamp+ value columns converted tostart/freq(inferred) /target. KernelSynth-1M fromtraining_corpus/kernel_synth_1m. - Spanish Weather: from
autogluon/chronos_datasets_extra(Kaggle-backed; needs Kaggle credentials), reshaped to 5 cities x {temp, pressure, humidity}. (Omitted unless built with Kaggle creds.) - Synthetic (
syn_*): regenerated with the open-sourceautoml/TempoPFNgenerators (Anomaly, ForecastPFN, GP, Sawtooth, SineWave, Spikes, Step, OU), 5,000 series each, fixed per-dataset seeds. Exact paper samples are not recoverable; these are statistically equivalent and reproducible from the documented seeds. - Paper-faithful downsampling (seeded):
buildings_900k-> 100k of ~1.8M series;weatherbench_daily-> 10k;era5_*-> 15 of 45 channels;cmip6_2000-> 22 of 53. - Cleaning: float32; inf -> NaN (missing values preserved as NaN); empty/all-NaN series dropped; frequency aliases normalized.
- Validation: #series and #channels checked against Table 5 (hard); max_length is informational (current upstream snapshots sometimes have longer raw series than Table 5).
Known deviations from the paper
- `syn_gp` uses length 2048 instead of 10,000. Exact GP at length 10,000 is computationally infeasible (Cholesky fails ->
symeigon 10k x 10k matrices, ~minutes each; ~24 h for 5,000 series). 2048 is a standard GP/KernelSynth prior length. - A few datasets (e.g. BDG-2, windfarmshourly) have longer max_length than Table 5 because the current upstream snapshot contains longer raw series. Values are unmodified.
Licensing
Derivative reconstruction; each subset is governed by its upstream license:
- LOTSA subsets: see
Salesforce/lotsa_data(per-dataset licenses). - Chronos subsets: see
autogluon/chronos_datasets. - Spanish Weather: Kaggle "energy-consumption-generation-prices-and-weather".
- Synthetic
syn_*: generated withautoml/TempoPFN(Apache-2.0).
Users must comply with the original licenses and cite the original dataset authors.
Citation
If you use this corpus, please cite both this dataset and the upstream sources.
This dataset (the reconstruction):
@misc{jua2026tsiclcorpus,
title = {TS-ICL Pretraining Corpus (reconstruction)},
author = {Jua},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus}},
note = {Community reconstruction of the TS-ICL (arXiv:2606.05878) pretraining mix}
}Upstream sources:
@article{lenaour2026tsicl,
title = {TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning},
author = {Le Naour, Etienne and Nabil, Tahar and Petralia, Adrien},
journal= {arXiv preprint arXiv:2606.05878},
year = {2026}
}
@article{woo2024moirai, title={Unified Training of Universal Time Series Forecasting Transformers (LOTSA)}, author={Woo, Gerald and others}, year={2024}}
@article{ansari2024chronos, title={Chronos: Learning the Language of Time Series}, author={Ansari, Abdul Fatir and others}, journal={arXiv:2403.07815}, year={2024}}
@misc{moroshan2025tempopfn, title={TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-Shot Time Series Forecasting}, author={Moroshan, Vladyslav and Siems, Julien and Zela, Arber and Carstensen, Timur and Hutter, Frank}, eprint={2510.25502}, year={2025}}