datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MA_Query_Expansion_MLT26ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.ML-Proto-Datasetnaver-news-summarization-ko
Naver-News-KO: A Korean News Summarization Dataset
A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from
Naver News over a ten-day window in July 2022. It was originally built for a
Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023.
A technical report documenting the collection protocol, corpus statistics, contamination analysis, and
reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.CCN
Towards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Route Features
Used to describe each route, including static features, dynamic features, and trajectory statistical features
N * 62
The estimated time of arrival for the routeThe total distance length… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/CCN.Taming-Hallucinations
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
CVPR 2026 Findings
Project Page | Paper | Code
Dataset Summary
This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation.
Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns
real videos into… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Taming-Hallucinations.NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
AoA range: −5° to +5°
Reynolds number range: 100,000 – 500,000
129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.TransitLM
TransitLM: Dataset Release & Evaluation Protocol
Dataset Description
TransitLM is a dataset for public transit route planning in Chinese urban environments, designed to support training and evaluation of language models that generate structured transit routes from origin-destination information. The full dataset covers four cities: Beijing, Shanghai, Shenzhen, and Chengdu, and includes coordinates, station sequences, transfer structure, line information, and route… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/TransitLM.faang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.MLMA_hate_speech
Disclaimer
This is a hate speech dataset (in Arabic, French, and English).
Offensive content that does not reflect the opinions of the authors.
Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis)
For more details about our dataset, please check our paper:
@inproceedings{ousidhoum-etal-multilingual-hate-speech-2019,
title = "Multilingual and Multi-Aspect Hate Speech Analysis",
author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.mlsum-it
Dataset Card for mlsum-it
Dataset Summary
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
More informations on the official dataset page HuggingFace page.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.SCASRec
SCASRec: A Self-Correcting and Auto-Stopping Model for Generative Route List Recommendation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Route Features
Used to describe each route, including static features, dynamic features, and trajectory statistical features
N * 62
The estimated time of arrival for the routeThe total distance length of the… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/SCASRec.MobilityBench
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
Paper | GitHub
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MobilityBench.GenMRP
GenMRP: A Generative Multi-Route Planning Framework for Efficient and Personalized Real-Time Industrial Navigation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Link Features
Includes the road segment attributes
K * 2 * N
Link lengthLink Lane width
Frequency Features
Logs the user's travel history within the past three months
K * 2 * 10 * 7
Delta… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/GenMRP.ml-intern-kpis
bep40/ml-intern-kpis
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset("bep40/ml-intern-kpis")
ASAP_OpenADMET_MLM
ASAP-OpenADMET MLM
ASAP_OpenADMET_MLM dataset from the ASAP Discovery-OpenADMET Antiviral Drug Discovery Challenge [1] [2] [3]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MLM (mouse liver microsomal intrinsic clearance in uL/min/mg) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
425
Recommended splittime
Recommended metric
MAE
References
[1]
ASAP Discovery
"ASAP… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ASAP_OpenADMET_MLM.ML-KEM-SideChannel-Traces
ML-KEM Side Channel Traces
Dataset Description
This dataset contains power traces captured from the Post-Quantum Cryptography (PQC) ML-KEM implementation of the PQM4[1] library (commit: a24bb4b), running on an STM32 Nucleo-L4R5ZI development board equipped with an ARM Cortex-M4 processor.
The traces were collected using a Rohde & Schwarz RTC1002 100 MHz digital oscilloscope. The purpose of this dataset is to evaluate the ML-KEM implementation for side-channel… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/ML-KEM-SideChannel-Traces.coconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.energy_datasetmlb-statcast-battersml-interview-examples-movielens-1mExpansionRx_OpenADMET_MLM_CLint
ExpansionRx-OpenADMET MLM CLint
MLM CLint dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MLM CLint of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
5692
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_MLM_CLint.genz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.Rash-Driving-Detection-on-Bikes-for-ML
Dataset Card for Rash Driving Detection on Bikes Using Mobile and Sensor Data
This dataset is designed to aid the detection of rash driving behavior on bikes using data collected from mobile and sensor-based systems. It includes sensor readings such as accelerometer values, orientation (azimuth, pitch, roll), and speed, with labels indicating whether the riding behavior is classified as rash or not.
Dataset Details
Dataset Description
This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Rash-Driving-Detection-on-Bikes-for-ML.cwe-workshop-datasetlotte_pooled_dev_search_mlateon
lotte_pooled_dev_search_mlateon
Multi-vector (late-interaction) embeddings of LoTTE pooled/dev/search (lotte/pooled/dev/search), encoded with
lightonai/mLateOn at revision edd378f99593c0ac8a15518b97ad89786b02685e.
Source data: ir_datasets lotte/pooled/dev/search (ir_datasets 0.6.3), which downloads lotte.tar.gz (md5 3b2e88b1d66933627462950b4c3f5d0f). The ColBERTv2 authors also publish LoTTE on the Hub as colbertv2/lotte, whose card gives this dataset's license; the data here was… See the full description on the dataset page: https://huggingface.co/datasets/robro612/lotte_pooled_dev_search_mlateon.trec-covid_mlateon
trec-covid_mlateon
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
lightonai/mLateOn at revision edd378f99593c0ac8a15518b97ad89786b02685e.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_mlateon.malayalam_news_classificationmlit-japan-real-estate-price-2005-2024q3
