datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cbb26-timeseries-db
CBB26 Timescale market data (cbb26-timeseries-db)
Public Timescale/Postgres replay shards for the cbb26 monorepo: canonical Coinbase Advanced Trade level-2 order book history stored in the market_data schema, packaged as restorable pg_dump files for research, corpus materialization, and reproducibility.
Canonical Hub repo: deusmos/cbb26-timeseries-db
Source-of-truth for this card (edit here, then publish): docs/datasets/cbb26-timeseries-db/README.md in the cbb26 git repo.… See the full description on the dataset page: https://huggingface.co/datasets/deusmos/cbb26-timeseries-db.Time-Series-Library
Time-Series-Library (TSLib)
TSLib is an open-source library for deep learning researchers, especially for deep time series analysis.
We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification.
This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/thuml/Time-Series-Library.TimeLens-100K
TimeLens-100K
📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data
✨ Dataset Description
TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro.
📊 Dataset Statistics
Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.precision-seed-germination-time-lapsePolymarket-v1
Polymarket-v1
A large-scale dataset of on-chain event logs from Polymarket v1, the prediction-market platform on the Polygon network. The repository contains four layers covering the full contract lifecycle from 2022-11-21 to 2026-04-28 (from first settlement to natural termination): OrderFilled/ (the raw on-chain trade tape), daily_aligned/ (the cleaned, metadata-enriched, and normalized analysis layer for Standard Binary markets, neg_risk=false), daily_aligned_multi/ (the… See the full description on the dataset page: https://huggingface.co/datasets/TimeSeventeen/Polymarket-v1.time-lapse-artifacts
Time-Lapse Artifacts
1,234 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through October 5, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,332,970,525,188 indexed video bytes
(approximately 2.33 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.Polymarket-v2
Polymarket-v2
A large-scale dataset of on-chain event logs from Polymarket v2, the prediction-market platform on the Polygon network. The repository contains three layers covering the full contract lifecycle from Polymarket v2 server start: OrderFilled/ (the raw on-chain trade tape), daily_aligned/ (the cleaned, metadata-enriched, and normalized analysis layer, Split by UTC). update daliy.
android-times-articles
Android Times — Articles Dataset Archive
Private dataset containing synthesized and processed article archives, multi-language transcripts, metadata, and editorial assets for Android Times.
Dataset Structure
articles/
├── en-US/ # English (United States) localized articles & scripts
├── ja-JP/ # Japanese localized articles & scripts
├── en-AU/ # Australian localized articles
├── en-CN/ # China localized English… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/android-times-articles.Time-300B
Dataset Card for Time-300B
This repository contains the Time-300B dataset of the paper Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts.
For details on how to use this dataset, please visit our GitHub page.
TimeLens2-93K
TimeLens2-93K
TimeLens2-93K is a large-scale, long-video temporal grounding dataset. This release contains 23,793 videos and 93,232 text–temporal interval pairs, including 12,091 multi-span pairs. The videos range from short clips to nearly 100 minutes and cover broad web domains such as entertainment, education, sports, news, science and technology, gaming, travel, vehicles, music, and daily life.
TimeLens2-93K offers a rare combination of scale, long-context coverage… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/TimeLens2-93K.time-series-datasetsplitted_PretrainGiftEvalai-timeline
AI Timeline Dataset
An open, dated, source-linked record of what artificial intelligence actually
did between July 2025 and today. Every row is a single real-world event with a
primary source attached.
3,478 events · 936 distinct publishers · 2025-07-01 to 2026-10-10
Maintained by Present of AI, a daily AI news site.
Updated as the timeline grows.
Why this exists
Most AI datasets are benchmarks or model outputs. This one is a record of
events: deployments, funding… See the full description on the dataset page: https://huggingface.co/datasets/presentofai/ai-timeline.timewarp
Timewarp datasets
This dataset contains molecular dynamics simulation data that was used to train the neural networks in the NeurIPS 2023 paper Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics by Leon Klein, Andrew Y. K. Foong, Tor Erlend Fjelde, Bruno Mlodozeniec, Marc Brockschmidt, Sebastian Nowozin, Frank Noé, and Ryota Tomioka.
Please see the accompanying GitHub repository.
This dataset consists of many molecular dynamics trajectories… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/timewarp.TimeChat-Online-139K
TimeChat-Online-139K
Project Page | Paper | GitHub | Model Checkpoint
⚠️ Important Notice: Research-Only Use
Before downloading or using this dataset, you must agree to the LICENSE terms.
This dataset contains video data that may be copyright-sensitive.
It is provided solely for non-commercial research and educational purposes.
By accessing this dataset, you confirm that you understand and agree to the terms in the LICENSE file.
📦 Dataset Overview
For flexible… See the full description on the dataset page: https://huggingface.co/datasets/yaolily/TimeChat-Online-139K.HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.TimeLens-Bench
TimeLens-Bench
📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data | 🏆 TimeLens-Bench Leaderboard
✨ Dataset Description
TimeLens-Bench is a comprehensive, high-quality evaluation benchmark for video temporal grounding, proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs.
During our annotation process, we identified critical quality issues within existing datasets and performed extensive manual corrections. We observed a… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-Bench.timememory
TimeMemory — private academic research artifacts
This repository is a private migration archive, not a public dataset release or a single uniform load_dataset table.
Companion code, detailed recovery manual and sanitized research/conversation history:
GitHub liofoil/timememory.
Start with GitHub RESTORE.md, CODEX_HANDOFF.md, and docs/RESEARCH_HISTORY.md.
Download only the explicit files in migration/artifacts_manifest.json, using the frozen HF commit recorded in the companion… See the full description on the dataset page: https://huggingface.co/datasets/liofoil/timememory.TimeBlind
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
Baiqi Li1, Kangyi Zhao2, Ce Zhang1, Chancharik Mitra3, Jean de Dieu Nyandwi3, Gedas Bertasius1 1University of North Carolina at Chapel Hill 2University of Pittsburgh 3Carnegie Mellon University
🏠Home Page | 🤗HuggingFace | 📖Paper | 🖥️ Code
Setup
git clone https://github.com/Baiqi-Li/TimeBlind.git
cd TimeBlind
git clone https://huggingface.co/datasets/BaiqiL/TimeBlind
Data… See the full description on the dataset page: https://huggingface.co/datasets/BaiqiL/TimeBlind.Timechat-OmniCaptioner-42K
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
🌟 Overview
TimeChat-Captioner is a multimodal model designed to generate detailed, time-aware, and structurally coherent captions for multi-scene videos. It effectively coordinates visual and audio information to provide comprehensive video descriptions.
🌐 Project Page: timechat-captioner.github.io
🏠 Model: TimeChat-Captioner (7B)
📚 Train Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/yaolily/Timechat-OmniCaptioner-42K.Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Y123-wed/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.Timeseries-PILE
Time Series PILE
The Time-series Pile is a large collection of publicly available data from diverse domains, ranging from healthcare to engineering and finance. It comprises of over 5
public time-series databases, from several diverse domains for time series foundation model pre-training and evaluation.
Time Series PILE Description
We compiled a large collection of publicly available datasets from diverse domains into the Time Series Pile. It has 13 unique domains of data… See the full description on the dataset page: https://huggingface.co/datasets/AutonLab/Timeseries-PILE.TIME-OutputThis repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for every experiment.
Note: These files are for building leaderboard and visualization; users do not need to download this directory.
features/: Statistical Features (tsfeatures)
Each dataset's features are saved to: output/features/{dataset}/{freq}/.
This directory stores the computed tsfeatures for the variates in the dataset. The folder contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Real-TSF/TIME-Output.pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
Pseudolabel Malaysian Youtube using Whisper Large V3 including Timestamp
how to prepare the dataset
wget https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp/resolve/main/prepared-pseudolabel.jsonl
huggingface-cli download --repo-type dataset \
--include 'output-audio-*.zip' \
--local-dir './' \
--max-workers 20 \
mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp.timeseries_dataPaderborn_Bearing_Run-to-Failure_Time-VaryingTime-Series-Forecasting-Benchmark-Datasets
Time Series Forecasting Benchmark Datasets
Documentation Language
简体中文 | English | Tiếng Việt
Dataset Download
https://huggingface.co/datasets/Duyu/Time-Series-Forecasting-Benchmark-Datasets/tree/main
https://github.com/duyu09/TimeSeries-Forecasting-Dataset/releases/download/v1.0.0/dataset.7z
Dataset Desc.
ETT The Electricity Transformer Temperature (ETT) dataset serves as a critical benchmark for evaluating electric power forecasting. It… See the full description on the dataset page: https://huggingface.co/datasets/Duyu/Time-Series-Forecasting-Benchmark-Datasets.TreeSatAI-Time-Series
TreeSatAI-Time-Series
This dataset was introduced in the ECCV24 paper OmniSat.
Ahlswede et al. (https://essd.copernicus.org/articles/15/681/2023/) introduced the TreeSatAI Benchmark Archive, a new dataset for tree species classification in Central Europe based on multi-sensor data from aerial,
Sentinel-1 and Sentinel-2. The dataset contains labels of 20 European tree species (i.e., 15 tree genera) derived from forest administration data of the federal state of Lower Saxony… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/TreeSatAI-Time-Series.Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Wenyan0110/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.synthetic-timeseries-data
cruscy data — evaluation sample
Three full days of real crypto market microstructure (Binance spot), prepared
for public evaluation: absolute prices, dates, and the instrument are withheld —
the shape of the day (tick-by-tick relative price, normalized volumes, trade
side, book imbalance) is fully preserved.
The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book
metrics, derived features, regime labels) with SQL console, backtest runner and
MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.
