datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
time-series-forecasting-datasetsmkdir -p dataset/ETT-small
mkdir -p dataset/electricity
mkdir -p dataset/traffic
mkdir -p dataset/weather
(
cd dataset
# -t 5: Retry up to 5 times on failure
# -nc: Skip download if file already exists (no-clobber)
# -q: Run quietly to suppress long logs (remove if not needed)
COMMON_ARGS="-t 5 -nc"
URL_PREFIX="https://huggingface.co/datasets/pkr7098/time-series-forecasting-datasets/blob/main"
wget $COMMON_ARGS $URL_PREFIX/ETTh1.csv &
wget $COMMON_ARGS… See the full description on the dataset page: https://huggingface.co/datasets/pkr7098/time-series-forecasting-datasets.time_series_forecastingforecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models"
This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.forecastingDataset from "Approaching Human-Level Forecasting with Language Models"
This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.surya-bench-flare-forecasting
Full-disk Solar Flare Forecasting Dataset
Dataset Summary
This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index.
Two derived binary labels are included for forecasting tasks:
label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.store-sales-time-series-forecasting
taken from this Kaggle competition:
Dataset Description
In this competition, you will predict sales for the thousands of product families sold at Favorita stores located in Ecuador. The training data includes dates, store and product information, whether that item was being promoted, as well as the sales numbers. Additional files include supplementary information that may be useful in building your models.
File Descriptions and Data Field Information… See the full description on the dataset page: https://huggingface.co/datasets/mrcksggcfc/store-sales-time-series-forecasting.multimodal-time-series-forecastinglaliga-football-match-forecasting
LaLiga Football Match Forecasting Dataset
Version 0.2.0 is a reproducible, match-level starting point for
time-aware LaLiga forecasting. It contains historical full-time results and
basic match statistics from Football-Data.co.uk, plus features calculated only
from matches that occurred before each target match.
Contents
pre_match_forecasting.parquet — the main Hugging Face dataset with
chronological train, validation, and test splits.… See the full description on the dataset page: https://huggingface.co/datasets/NewSonnet/laliga-football-match-forecasting.m5-retail-demand-forecasting-benchmarks
M5 Retail Demand Forecasting & Inventory Risk Benchmarks
This dataset contains the heavily processed artifacts, extracted time-series features, baseline benchmarks, and model artifacts for the M5 Retail Demand Forecasting dataset.
It includes:
Over 1GB of highly engineered temporal, pricing, and calendar features.
Volatility and shortfall risk metrics for 42,840 time series.
XGBoost, Prophet, and SARIMA predictions (point + 95% intervals).
Isolation Forest anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/snchakri/m5-retail-demand-forecasting-benchmarks.Electricity_Load_Forecasting_using_LLMSgoes-omni-electron-flux-forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/goes-omni-electron-flux-forecasting.m5-sales-forecasting-datasetssuperkart-sales-forecastingweather-forecasting-challenge
Dataset Description
Data Overview
The WiDS Datathon 2023 focuses on a prediction task involving forecasting sub-seasonal temperatures (temperatures over a two-week period, in our case) within the United States. We are using a pre-prepared dataset consisting of weather and climate information for a number of US locations, for a number of start dates for the two-week observation, as well as the forecasted temperature and precipitation from a number of weather… See the full description on the dataset page: https://huggingface.co/datasets/serenia-science/weather-forecasting-challenge.restaurant-inventory-forecasting
🍽️ Restaurant Inventory Forecasting Dataset
Overview
High-quality synthetic training data for post-training a small LLM (Qwen2.5-1.5B) to become an expert reasoning engine for restaurant inventory forecasting.
Dataset Structure
Split
Examples
Purpose
sft_train
~23,750
SFT training (instruction-response pairs with CoT reasoning)
sft_eval
~1,250
SFT evaluation
grpo_train
~4,500
GRPO RL training (prompts + ground truth solutions)
grpo_eval
~500… See the full description on the dataset page: https://huggingface.co/datasets/EpicVic2193/restaurant-inventory-forecasting.market-forecasting-dataset
Crypto Market Daily Reviews & Sentiment Signals
A dataset for backtesting whether the news background of daily crypto market reviews has predictive power over subsequent price action.
It is built from the archive of daily cryptocurrency market reviews published on eraperemen.info (June 2023 — October 2026, ~700 reviews). Each review was converted into structured market-sentiment signals for Bitcoin and altcoins by an LLM pipeline, producing a machine-readable time series of… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/market-forecasting-dataset.en-forecasting-bigdata-sftgrocery-sales-forecastingbiomedical-forecasting-lightningrod
Biomedical Forecasting Dataset
A dataset of 1444 binary forecasting questions about biomedical and public health outcomes. Each question is a forward-looking prediction (Yes/No) about a real event, grounded in news and labeled with the actual outcome.
What is in this dataset?
Questions: FDA drug approvals, clinical trial results (Phase 2/3), WHO and CDC declarations, vaccine development, disease outbreaks, gene therapy, and public health policy.
Grounded in real news:… See the full description on the dataset page: https://huggingface.co/datasets/Ainoafv/biomedical-forecasting-lightningrod.africa-synth-retail-and-ecommerce-demand-forecasting-datasets-nigeria
Demand Forecasting Datasets | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-demand-forecasting-datasets-nigeria.ts-forecasting-benchmarknigerian_energy_and_utilities_demand_forecasting
Nigerian Energy & Utilities – Demand Forecasting | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_energy_and_utilities_demand_forecasting.Store-Sales-Time-Series-Forecasting-result-0.45607en-forecasting-bigdata-query-parsedfreeform-forecasting
Freeform Forecasting Dataset
Dataset for free-form forecasting questions generated from news articles, designed to evaluate AI models' ability to make predictions about future events.
Dataset Overview
This dataset contains 71,389 forecasting questions across three splits:
Train: 70,185 questions
Validation: 204 questions
Test: 1,000 questions
Dataset Structure
Fields Description
Field
Type
Description
question_title
string
The main… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/freeform-forecasting.inductive-forecasting-data
Inductive Forecasting Study — Anonymous Data Release
This repository is the anonymous data companion to a paper studying behavioral
signatures of inductive reasoning in language-model forecasts. It packages the
frozen inputs, model responses, row-level scores, and aggregate result artifacts
used by the paper's four main experiments, together with synthetic appendix
transfer studies.
The release is organized as Hugging Face dataset configurations so each study can
be loaded… See the full description on the dataset page: https://huggingface.co/datasets/od2961/inductive-forecasting-data.energy-forecasting-filesForecastingv1V1 of a custom event forecasting dataset.
Data is not very well filtered or well formatted.
wwtd-forecasting-demonigerian_transport_and_logistics_inventory_forecasting
Nigeria Transport & Logistics – Inventory Forecasting | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_transport_and_logistics_inventory_forecasting.
