Team Ai
Datasetpublic

JuaAI/ts-icl-pretraining-corpus

TS-ICL Pretraining Corpus (community reconstruction) A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and… See the full description on the dataset page: https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes1.2kdownloads
Dataset Card

TS-ICL Pretraining Corpus (community reconstruction)

A unified, cleaned reconstruction of the univariate pretraining corpus described in Table 5 of TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning (Le Naour, Nabil & Petralia, EDF R&D; arXiv:2606.05878). The TS-ICL authors did not release their pretraining data pipeline, so this corpus is rebuilt from the named upstream sources (LOTSA, Chronos, and the TempoPFN synthetic generators) and normalized into a single schema.

Curated by: Jua — reconstruction, normalization, synthetic generation, validation and packaging.

Not an official release. This is a derivative dataset built by Jua for research reproducibility; it is not affiliated with or endorsed by the TS-ICL authors. Each subset retains the license of its upstream source (see Licensing).

At a glance

  • —39 datasets, 1,399,719 series, 8,690,897,054 observations
  • —Sources: LOTSA (21), Chronos (10), TempoPFN synthetic (8)
  • —Single consistent schema, sharded Parquet (zstd)

Repository layout

data/<config>/*.parquet     # one folder per dataset (= HF config)
meta/datasets.yaml          # build registry (sources, fingerprints, subsampling)
meta/reports/<config>.json  # per-dataset build/QA report
meta/qa.json                # aggregated QA
QUALITY_REPORT.md           # human-readable QA table

Schema

Every row is one (possibly multivariate) time series:

fieldtypedescription
item_idstringunique id within a dataset
datasetstringdataset name (= config name)
domainstringEnergy / Climate / Traffic / Cloud / Web / Econ-Fin / Health / Synthetic
sourcestringlotsa / chronos / chronos_extra / tempopfn
freqstringpandas offset alias (null for synthetic)
starttimestamp[s]start time (null for synthetic)
targetlist<list<float32>>values as [num_channels, length] (univariate = [1, length])
num_channelsint32number of channels
lengthint32series length
weightfloat32Table-5 sampling coefficient (training-sampler metadata)

How to load

python
from datasets import load_dataset

# one dataset (config)
ds = load_dataset("<repo_id>", "nn5_weekly", split="train")

# everything, streamed
ds = load_dataset("<repo_id>", "all", split="train", streaming=True)

Datasets (provenance)

configpaper namesourceupstream locationdomainfreqweightseries
bdg2_bullBDG-2 BulllotsabullEnergyH2541
bdg2_foxBDG-2 Foxlotsabdg-2_foxEnergyH5135
bdg2_pantherBDG-2 Pantherlotsabdg-2_pantherEnergyH2.5105
buildings_900kBuildingsBench900klotsabuildings_900kEnergyH0.02048100,000
residential_load_powerResidential Load Powerlotsaresidential_load_powerEnergy1T1.2271
residential_pv_powerResidential PV Powerlotsaresidential_pv_powerEnergy1T1.5233
china_air_qualityChina Air Qualitylotsachina_air_qualityClimateH0.3437
cmip6_2000CMIP6 2000lotsacmip6_2000Climate6H0.0578,192
era5_1989ERA5 1989lotsaera5_1989ClimateH0.0858,192
era5_1990ERA5 1990lotsaera5_1990ClimateH0.0858,192
era5_1991ERA5 1991lotsaera5_1991ClimateH0.0858,192
subseasonalSubseasonallotsasubseasonalClimate1D0.3862
subseasonal_precipSubseasonal Precipitationlotsasubseasonal_precipClimate1D1.2862
pems04PEMS04lotsaPEMS04Traffic5T1.2307
pems07PEMS07lotsaPEMS07Traffic5T1.2883
pems08PEMS08lotsaPEMS08Traffic5T2.1170
q_trafficQ-TRAFFIClotsaQ-TRAFFICTraffic15T0.02445,148
alibaba_cluster_trace_2018Alibaba Cluster Trace 2018lotsaalibaba_cluster_trace_2018Cloud5T0.00958,409
monash_m3_monthlyMonash M3 Monthlylotsamonash_m3_monthlyEcon/FinM0.721,428
nn5_weeklyNN5 Weeklylotsann5_weeklyEcon/FinW5111
project_tychoProject Tycholotsaproject_tychoHealthW0.211,258
australian_electricityAustralian Electricitychronosmonash_australian_electricityEnergy30T2205
wind_farms_hourlyWind Farms Hchronoswind_farms_hourlyEnergyH4337
wind_farms_dailyWind Farms Dchronoswind_farms_dailyEnergyD2337
weatherbench_dailyWeatherbench dailychronosweatherbench_dailyClimate1D0.102410,000
mexico_city_bikesMexico City Bikeschronosmexico_city_bikesTrafficH2.5494
taxi_30minTaxi (30 Min.)chronostaxi_30minTraffic30T0.882,428
taxi_1hTaxi (Hourly)chronostaxi_1hTrafficH0.882,428
uber_tlc_hourlyUber TLC (Hourly)chronosuber_tlc_hourlyTrafficH4262
wiki_daily_100kWiki Dailychronoswiki_daily_100kWebD0.00512100,000
kernel_synth_1mKernel Synth 1Mchronostraining_corpus/kernel_synth_1mSyntheticNone0.0010241,000,000
syn_anomalyAnomalytempopfnAnomalyGeneratorSyntheticNone0.02565,000
syn_forecastpfnForecastPFNtempopfnForecastPFNGeneratorSyntheticNone15,000
syn_gpGPtempopfnGPGeneratorSyntheticNone0.40965,000
syn_sawtoothSawtoothtempopfnSawToothGeneratorSyntheticNone0.05125,000
syn_sinewaveSinewavetempopfnSineWaveGeneratorSyntheticNone0.10245,000
syn_spikesSpikestempopfnSpikesGeneratorSyntheticNone0.02565,000
syn_stepSteptempopfnStepGeneratorSyntheticNone0.05125,000
syn_ouOUtempopfnOrnsteinUhlenbeckProcessGeneratorSyntheticNone0.40965,000

Reconstruction methodology

  • —LOTSA subsets: from Salesforce/lotsa_data (Arrow); targets passed through.
  • —Chronos subsets: from autogluon/chronos_datasets; per-series timestamp + value columns converted to start / freq (inferred) / target. KernelSynth-1M from training_corpus/kernel_synth_1m.
  • —Spanish Weather: from autogluon/chronos_datasets_extra (Kaggle-backed; needs Kaggle credentials), reshaped to 5 cities x {temp, pressure, humidity}. (Omitted unless built with Kaggle creds.)
  • —Synthetic (syn_*): regenerated with the open-source automl/TempoPFN generators (Anomaly, ForecastPFN, GP, Sawtooth, SineWave, Spikes, Step, OU), 5,000 series each, fixed per-dataset seeds. Exact paper samples are not recoverable; these are statistically equivalent and reproducible from the documented seeds.
  • —Paper-faithful downsampling (seeded): buildings_900k -> 100k of ~1.8M series; weatherbench_daily -> 10k; era5_* -> 15 of 45 channels; cmip6_2000 -> 22 of 53.
  • —Cleaning: float32; inf -> NaN (missing values preserved as NaN); empty/all-NaN series dropped; frequency aliases normalized.
  • —Validation: #series and #channels checked against Table 5 (hard); max_length is informational (current upstream snapshots sometimes have longer raw series than Table 5).

Known deviations from the paper

  • —`syn_gp` uses length 2048 instead of 10,000. Exact GP at length 10,000 is computationally infeasible (Cholesky fails -> symeig on 10k x 10k matrices, ~minutes each; ~24 h for 5,000 series). 2048 is a standard GP/KernelSynth prior length.
  • —A few datasets (e.g. BDG-2, windfarmshourly) have longer max_length than Table 5 because the current upstream snapshot contains longer raw series. Values are unmodified.

Licensing

Derivative reconstruction; each subset is governed by its upstream license:

  • —LOTSA subsets: see Salesforce/lotsa_data (per-dataset licenses).
  • —Chronos subsets: see autogluon/chronos_datasets.
  • —Spanish Weather: Kaggle "energy-consumption-generation-prices-and-weather".
  • —Synthetic syn_*: generated with automl/TempoPFN (Apache-2.0).

Users must comply with the original licenses and cite the original dataset authors.

Citation

If you use this corpus, please cite both this dataset and the upstream sources.

This dataset (the reconstruction):

bibtex
@misc{jua2026tsiclcorpus,
  title        = {TS-ICL Pretraining Corpus (reconstruction)},
  author       = {Jua},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/JuaAI/ts-icl-pretraining-corpus}},
  note         = {Community reconstruction of the TS-ICL (arXiv:2606.05878) pretraining mix}
}

Upstream sources:

bibtex
@article{lenaour2026tsicl,
  title  = {TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning},
  author = {Le Naour, Etienne and Nabil, Tahar and Petralia, Adrien},
  journal= {arXiv preprint arXiv:2606.05878},
  year   = {2026}
}
@article{woo2024moirai, title={Unified Training of Universal Time Series Forecasting Transformers (LOTSA)}, author={Woo, Gerald and others}, year={2024}}
@article{ansari2024chronos, title={Chronos: Learning the Language of Time Series}, author={Ansari, Abdul Fatir and others}, journal={arXiv:2403.07815}, year={2024}}
@misc{moroshan2025tempopfn, title={TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-Shot Time Series Forecasting}, author={Moroshan, Vladyslav and Siems, Julien and Zela, Arber and Carstensen, Timur and Hutter, Frank}, eprint={2510.25502}, year={2025}}