datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_objectnetclimbmix-400b-shuffleclinical-trials-protocolsclinc_oos
Dataset Card for CLINC150
Dataset Summary
Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example. We introduce a new dataset that includes queries that are out-of-scope (OOS), i.e., queries that do not fall into any of the system's supported intents. This poses a new challenge because models cannot assume that every query at inference… See the full description on the dataset page: https://huggingface.co/datasets/clinc/clinc_oos.wds_imagenet_sketchClimSim_low-res-testClimateFEVER_test_top_250_only_w_correct-v2
ClimateFEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio.
Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video.
🚀 Explore it in LightlyStudio
hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips
cd epic-kitchens-100-clips
pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.ClimSim_high-resThe corresponding GitHub repo can be found here:https://github.com/leap-stc/ClimSim
Read more: https://arxiv.org/abs/2306.08754.
Nemotron-ClimbLab
ClimbLab Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.ClimSim_low-res-expanded-testwds_imagenet-rwds_imagenet-awds_imagenet1kClinicalAgentBenchMore detail about the dataset and the agentic framework can be found in https://github.com/BlueZeros/ReflecTool
ClimSim_low-res-expandedThis is an expanded version of ClimSim_low-res. Each '.mlexpand.' file contains the same variables as in the corresponding '.mli.' file in ClimSim_low-res but also includes additional variables such as dynamical forcing, convection memory, cos/sin of latitude.
Read more about these expanded input features at Section 6.3.3 in the SI of "ClimSim-Online: A Large Multi-scale Dataset and Framework for Hybrid ML-physics Climate Emulation": https://arxiv.org/abs/2306.08754.
wds_imagenetv2Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.hdr-demo-clips
HDR Demo Clips (Lightricks SDR→HDR)
Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2).
Each clip contains:
hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred)
sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred)
thumbnail.jpg — 280px preview from the middle frame
Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.beir-nl-cqadupstack
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.clientsinstructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.repo2rlenv-cli-gym
Repo2RLEnv CLI-Gym
Start with a healthy repository, synthesize a disruption and recovery, verify that damage breaks the original tests and recovery restores them, then export an environment-repair instruction and a deterministic verifier.
Contains 25 Harbor tasks generated with the owned
cli_gym recipe in Repo2RLEnv.
Browse the complete task bundles in Harbor Visualiser or
open the task folders. Each folder is a runnable Harbor task:
tasks/<task_id>/
├── task.toml… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-cli-gym.ClimbMix
ClimbMix
About
🧗 A more convenient ClimbMix (https://arxiv.org/abs/2504.13161)
Description
Unfortunately, the original ClimbMix (https://huggingface.co/datasets/nvidia/ClimbMix) has four main inconveniences:
It is in GPT2 tokens, meaning you have to detokenize it to inspect it or use it with another tokenizer.
It contains all of the 20 clusters in order together (in the same "subset"), so you have to load the whole dataset in memory (~1TB) and shuffle it… See the full description on the dataset page: https://huggingface.co/datasets/gvlassis/ClimbMix.wds_fer2013CLIcK
CLIcK 🇰🇷🧠
A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
Introduction 🎉
CLIcK (Cultural and Linguistic Intelligence in Korean) is a comprehensive dataset designed to evaluate cultural and linguistic intelligence in the context of Korean language models. In an era where diverse language models are continually emerging, there is a pressing need for robust evaluation datasets, especially for non-English languages like Korean. CLIcK… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/CLIcK.code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.BabyLMDataset for the shared baby language modeling task.
The goal is to train a language model from scratch on this data which represents
roughly the amount of text and speech data a young child observes.cobotmagic_Sim_click_bell
