Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips 37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio. Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video. 🚀 Explore it in LightlyStudio hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips cd epic-kitchens-100-clips pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K3 likes16k downloads2d agoHugging Face02nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B130 likes7.4k downloads1y agoHugging Face03KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes2.7k downloads14d agoHugging Face04OptimalScale /ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper. We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.tabulartext-generation100M<n<1B37 likes2.5k downloads1y agoHugging Face05scikit-fingerprints /MoleculeNet_ClinTox MoleculeNet ClinTox Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary. Characteristic Description Tasks 2 Task type multitask classification Total samples 1477 Recommended split scaffold Recommended metric AUROC References [1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.tabulartabular-classification1K<n<10K0 likes2.2k downloads2y agoHugging Face06tjhunter /climate-tracetabular100M<n<1B2 likes1.6k downloads2y agoHugging Face07ClimDetect /ClimDetect ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution Details ClimDetect combines daily 2m surface temperature (‘tas’), 2m specific humidity (‘huss’) and total precipitation (‘pr’) from CMIP6 historical and ScenarioMIP (ssp245 and ssp370) experiments (Eyring et al. 2016*). See Clim_Detect_ds_info.csv for CMIP6 models and their data references. From the original CMIP6 model output, daily snapshots are chosen, sub-sampled every 5 days using… See the full description on the dataset page: https://huggingface.co/datasets/ClimDetect/ClimDetect.tabular1M<n<10M3 likes1.6k downloads1y agoHugging Face08scikit-learn /credit-card-clients Default of Credit Card Clients Dataset The following was retrieved from UCI machine learning repository. Dataset Information This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005. Content There are 25 variables: ID: ID of each client LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.tabular10K<n<100K10 likes1.1k downloads4y agoHugging Face09clips /mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.tabularquestion-answering10M<n<100M37 likes1k downloads4y agoHugging Face10Greynar /ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 101, "total_frames": 53052, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:101" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOneRedTriangle.tabularrobotics10K<n<100K0 likes856 downloads14d agoHugging Face11Greynar /ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleMultiplePieceThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 150, "total_frames": 79168, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:150" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleMultiplePiece.tabularrobotics10K<n<100K0 likes679 downloads14d agoHugging Face12DUDE-Framework /Real-UI-Clickboxes RUC: Real UI Clickboxes Click carefully, even when the page is trying to trick you! 👀 Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents. ACL Anthology: https://aclanthology.org/2026.acl-long.310/ PDF: https://aclanthology.org/2026.acl-long.310.pdf DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.imageimage-text-to-text1K<n<10K1 likes670 downloads3mo agoHugging Face13DataForGood /climate-misinformation-RCoTtabular1K<n<10K1 likes657 downloads5mo agoHugging Face14setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K1 likes614 downloads6mo agoHugging Face15Changyeli03 /AA_preference_cliptabular10K<n<100K0 likes563 downloads2y agoHugging Face16hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes496 downloads1y agoHugging Face17BeIR /climate-fever-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.tabulartext-retrieval1K<n<10K0 likes480 downloads4y agoHugging Face18ivannatarov /clickhouse_coursetabular100M<n<1B1 likes466 downloads5mo agoHugging Face19WayenVan /how2sign-front-clips How2Sign RGB Front Clips (VideoFolder) This directory packages the frontal-view, sentence-level How2Sign RGB clips in Hugging Face VideoFolder format. It contains the official train, validation, and test splits with English sentence annotations. Directory layout how2sign_videofolder/ ├── README.md ├── train/ │ ├── metadata.parquet │ └── shard_001_032/ ... shard_032_032/ ├── validation/ │ ├── metadata.parquet │ └── shard_001_018/ ... shard_018_018/ └──… See the full description on the dataset page: https://huggingface.co/datasets/WayenVan/how2sign-front-clips.tabulartranslation10K<n<100K0 likes463 downloads17d agoHugging Face20hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes458 downloads2y agoHugging Face21closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes441 downloads4y agoHugging Face22ching-goodfire /MAPS-ClinVar-VKS-Embeddings-L80 MAPS ClinVar/VKS ESM-C layer-80 difference fields Mutant-minus-wild-type difference fields at block 80 of ESM-C 6B for all 200,913 human missense variants of known clinical significance in the MAPS ClinVar/VKS set: the complete 12,565-variant held-out test split and the complete 188,348-variant training pool, no sampling on either side. 260.8 GB of raw fp16 payload, 154.0 GB on disk in 262 shards, one parquet row per variant, every row self-describing — no join with any other… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80.tabular100K<n<1M0 likes438 downloads2mo agoHugging Face23RoboCOIN /AIRBOT_MMK2_storage_remote_control_clip_box_water_bottlegated AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle.tabularrobotics1K<n<10K0 likes434 downloads10mo agoHugging Face24louisbrulenaudet /clinical-trials Clinical Trials Dataset A comprehensive dataset of clinical trials sourced from ClinicalTrials.gov, featuring structured metadata, detailed study information, and pre-computed semantic embeddings for machine learning applications in biomedical research. Dataset Description This dataset provides a rich collection of clinical trial information systematically collected from the official ClinicalTrials.gov database. Each record contains detailed study metadata, eligibility… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/clinical-trials.tabularquestion-answering100K<n<1M41 likes414 downloads1y agoHugging Face25anamnesis-data /clinical_trials_history_sample Clinical Trials Version History: Sample A free sample of the Clinical Trials Version History dataset by Anamnesis Data: structured medical and scientific data for biotech, pharma and healthcare research. Public trial registries show only a trial's latest state. This dataset keeps every version, so you can see what a trial said on any past date, or when a sponsor moved a completion date. The sample holds 20 complete trials (741 versions), one Parquet file per table (79 tables… See the full description on the dataset page: https://huggingface.co/datasets/anamnesis-data/clinical_trials_history_sample.tabulartabular-classification100K<n<1M2 likes413 downloads10d agoHugging Face26lightonai /climate-fever-decontaminated climate-fever (Decontaminated) A decontaminated version of the climate-fever dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/climate-fever-decontaminated.tabulartext-retrieval1M<n<10M0 likes405 downloads7mo agoHugging Face27cm3r /meteo-france-climatologie-quotidienne Climatologie quotidienne Météo-France Archive des observations climatologiques quotidiennes des stations Météo-France, de 1786-05-01 à 2026-09-10, en Parquet partitionné par département et par période. Redistribution non officielle. Ce dépôt n'est pas publié par Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0. Source et attribution Source : Météo-France, données publiques de climatologie de base. Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.tabular100M<n<1B0 likes381 downloads26d agoHugging Face28closji /flickr30k_clip-ViT-B-32-caption_pairstabular10M<n<100M4 likes323 downloads4y agoHugging Face29rabuahmad /climatecheck The ClimateCheck Dataset This dataset is used for the ClimateCheck: Scientific Fact-checking of Social Media Posts on Climate Change Shared Task. The 2025 iteration was hosted at the Scholarly Document Processing workshop at ACL 2025, and a new 2026 iteration will be hosted at the Natural Scientific Language Processing workshop at LREC 2026. 2026 Update For running the next iteration of the task, we added manually labelled training data, resulting in 3023… See the full description on the dataset page: https://huggingface.co/datasets/rabuahmad/climatecheck.tabulartext-retrieval1K<n<10K8 likes281 downloads4mo agoHugging Face30ranbyDipz /sih-lidar-cliptabularn<1K0 likes262 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.