datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio.
Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video.
🚀 Explore it in LightlyStudio
hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips
cd epic-kitchens-100-clips
pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.climate-traceClimDetect
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
Details
ClimDetect combines daily 2m surface temperature (‘tas’), 2m specific humidity (‘huss’) and total precipitation (‘pr’) from CMIP6 historical and ScenarioMIP (ssp245 and ssp370) experiments (Eyring et al. 2016*).
See Clim_Detect_ds_info.csv for CMIP6 models and their data references. From the original CMIP6 model output, daily snapshots are chosen, sub-sampled every 5 days using… See the full description on the dataset page: https://huggingface.co/datasets/ClimDetect/ClimDetect.credit-card-clients
Default of Credit Card Clients Dataset
The following was retrieved from UCI machine learning repository.
Dataset Information
This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005.
Content
There are 25 variables:
ID: ID of each client
LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit
SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 101,
"total_frames": 53052,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOneRedTriangle.ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleMultiplePieceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 150,
"total_frames": 79168,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleMultiplePiece.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.climate-misinformation-RCoTclimbing-holds
[!IMPORTANT]
This dataset is in construction. The current files are raw scans intended for establishing the structure.
Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide.
GUI for contributions
https://setrsoft.github.io/holds-dataset-hub/
Or send your files here
Climbing Holds 3D dataset (SetRsoft)
📋 Project Overview
This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.AA_preference_clipin1k_clip_qwen25vl_3b_224res_64tokens_new_ptclimate-fever-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.clickhouse_coursehow2sign-front-clips
How2Sign RGB Front Clips (VideoFolder)
This directory packages the frontal-view, sentence-level How2Sign RGB clips in
Hugging Face VideoFolder format. It contains the official train, validation,
and test splits with English sentence annotations.
Directory layout
how2sign_videofolder/
├── README.md
├── train/
│ ├── metadata.parquet
│ └── shard_001_032/ ... shard_032_032/
├── validation/
│ ├── metadata.parquet
│ └── shard_001_018/ ... shard_018_018/
└──… See the full description on the dataset page: https://huggingface.co/datasets/WayenVan/how2sign-front-clips.in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptcc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13MAPS-ClinVar-VKS-Embeddings-L80
MAPS ClinVar/VKS ESM-C layer-80 difference fields
Mutant-minus-wild-type difference fields at block 80 of ESM-C 6B for
all 200,913 human missense variants of known clinical
significance in the MAPS ClinVar/VKS set: the complete
12,565-variant held-out test split and the complete
188,348-variant training pool, no sampling on either side.
260.8 GB of raw fp16 payload, 154.0 GB on disk in 262 shards, one parquet row per variant, every row self-describing — no join with
any other… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80.AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle.clinical-trials
Clinical Trials Dataset
A comprehensive dataset of clinical trials sourced from ClinicalTrials.gov, featuring structured metadata, detailed study information, and pre-computed semantic embeddings for machine learning applications in biomedical research.
Dataset Description
This dataset provides a rich collection of clinical trial information systematically collected from the official ClinicalTrials.gov database. Each record contains detailed study metadata, eligibility… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/clinical-trials.clinical_trials_history_sample
Clinical Trials Version History: Sample
A free sample of the Clinical Trials Version History dataset by
Anamnesis Data: structured medical and scientific data for biotech,
pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or
when a sponsor moved a completion date.
The sample holds 20 complete trials (741 versions), one Parquet file per table
(79 tables… See the full description on the dataset page: https://huggingface.co/datasets/anamnesis-data/clinical_trials_history_sample.climate-fever-decontaminated
climate-fever (Decontaminated)
A decontaminated version of the climate-fever dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/climate-fever-decontaminated.meteo-france-climatologie-quotidienne
Climatologie quotidienne Météo-France
Archive des observations climatologiques quotidiennes des stations
Météo-France, de 1786-05-01 à 2026-09-10,
en Parquet partitionné par département et par période.
Redistribution non officielle. Ce dépôt n'est pas publié par
Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0.
Source et attribution
Source : Météo-France, données publiques de climatologie de base.
Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.flickr30k_clip-ViT-B-32-caption_pairsclimatecheck
The ClimateCheck Dataset
This dataset is used for the ClimateCheck: Scientific Fact-checking of Social Media Posts on Climate Change Shared Task.
The 2025 iteration was hosted at the Scholarly Document Processing workshop at ACL 2025, and a new 2026 iteration will be hosted at the Natural Scientific Language Processing workshop at LREC 2026.
2026 Update
For running the next iteration of the task, we added manually labelled training data, resulting in 3023… See the full description on the dataset page: https://huggingface.co/datasets/rabuahmad/climatecheck.sih-lidar-clip
