datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio.
Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video.
🚀 Explore it in LightlyStudio
hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips
cd epic-kitchens-100-clips
pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.credit-card-clients
Default of Credit Card Clients Dataset
The following was retrieved from UCI machine learning repository.
Dataset Information
This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005.
Content
There are 25 variables:
ID: ID of each client
LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit
SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.clinc150MTS_Dialogue-Clinical_Note
MTS Dialogue (Clinical Note Summarisation)
Main Dataset
The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents).
The training set consists of 1,201 pairs of conversations and associated summaries.
The validation set consists of 100 pairs of conversations and their summaries.
The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptclimate-fever-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.showdown-clicks
showdown-clicks
General Agents
🤗 Dataset | GitHub
showdown is a suite of offline and online benchmarks for computer-use agents.
showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents.
As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.detector-clickbait-br-datasets
Detector Clickbait BR - Datasets
Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait.
📚 Descrição dos Datasets
1. detector-clickbait-br-raw.csv
Dataset original contendo os dados iniciais sem processamento.
Características:
Dados brutos coletados originalmente
Pode conter duplicatas
Pode conter valores nulos
Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.LDS-retrain-bank-adamw-N16k-bs256-clip1.0Climate-Change-Indicators-For-African-Countries
Climate Change Indicators for African Countries
This repository contains time-series datasets for key climate change indicators for 54 African countries. The data is sourced from The World Bank and has been cleaned, processed, and organized for analysis.
Each country has its own set of files, including a main CSV dataset and a corresponding datacard in Markdown format. The data covers the period from 1960 to 2024, where available.
Repository Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/LopsidedLion/Climate-Change-Indicators-For-African-Countries.Climate-Change-Indicators-For-African-Countries
Climate Change Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Climate-Change-Indicators-For-African-Countries.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.clickbait_title_classificationDataset introduced in Stop Clickbait: Detecting and Preventing Clickbaits in Online News Mediaby Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, Niloy Ganguly
Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. "Stop Clickbait: Detecting and Preventing Clickbaits in Online News Media”. In Proceedings of the 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), San Fransisco, US, August 2016.
Cite:… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/clickbait_title_classification.ExpansionRx_OpenADMET_RLM_CLint
ExpansionRx-OpenADMET RLM CLint
RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules.
Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint.
Characteristic
Description
Tasks
1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.climate-stance-detectionBiogen_ADME_HLM_CLint
Biogen ADME HLM CLint
Biogen_ADME_HLM_CLint dataset from the Biogen ADME benchmark [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict log10 of human liver microsomal intrinsic clearance (HLM CLint, in mL/min/kg) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3087
Recommended split
time
Recommended metricMAE
References
[1]
Fang, Cheng, et al.
"Prospective Validation of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/Biogen_ADME_HLM_CLint.ExpansionRx_OpenADMET_HLM_CLint
ExpansionRx-OpenADMET HLM CLint
HLM CLint dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict HLM CLint of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
4541
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_HLM_CLint.climate-news-articles
🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat
🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not.
🗺️ Le contexte
Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.Biogen_ADME_RLM_CLint
Biogen ADME RLM CLint
Biogen_ADME_RLM_CLint dataset from the Biogen ADME benchmark [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict log10 of rat liver microsomal intrinsic clearance (RLM CLint, in mL/min/kg) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3054
Recommended split
time
Recommended metric
MAE
References
[1]
Fang, Cheng, et al.
"Prospective Validation of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/Biogen_ADME_RLM_CLint.CLIP-ViT-L-14-336-L20-features
OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels
Feature activation max visualization of the 4096 Features @ L20
CLIP+BLIP labels (may or may not describe what a neuron truly encodes!)
⚠️ May contain sensitive images, albeit abstract. Use responsibly!
Examples:
ExpansionRx_OpenADMET_MLM_CLint
ExpansionRx-OpenADMET MLM CLint
MLM CLint dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MLM CLint of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
5692
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_MLM_CLint.clinical-trial-outcomes-2020plus
Clinical Trial Outcomes (2020+) with Normalized Endpoints
124,790 normalized endpoints across 14,170 clinical studies that started on or after
2020-01-01 and have posted results on ClinicalTrials.gov.
Snapshot: 2026-09-08. Source: ClinicalTrials.gov API v2 (U.S. National Library of Medicine).
Built with ctgov — the same normalizer, released as an
MIT-licensed package with zero dependencies. So this snapshot is not a dead artifact: you can
re-run it against the live registry, or… See the full description on the dataset page: https://huggingface.co/datasets/GooseWithStories/clinical-trial-outcomes-2020plus.csa-clinical-stage-asset-intelligence-sample
CSA — Clinical-Stage Asset Intelligence · Free Sample
Clinical trials, FDA, and SEC — linked to the drug asset and the listed sponsor, with a
forward catalyst calendar. This is a free 150-row sample of the nearest-term
catalysts; the full snapshot carries 2,221 forward catalysts (955 linked to
124 listed sponsors) and 1,890 resolved assets.
Data, not investment advice. CSA is information, not a recommendation to buy, sell,
or hold any security. Estimated catalyst dates (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/csa-clinical-stage-asset-intelligence-sample.ClinicalTrial-gov_QA
