Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clip-benchmark /wds_objectnetimage1K<n<10K4 likes73k downloads4y agoHugging Face02karpathy /climbmix-400b-shuffle73 likes47k downloads7mo agoHugging Face03Parexel /clinical-trials-protocolstext10K<n<100K4 likes37k downloads1mo agoHugging Face04clinc /clinc_oos Dataset Card for CLINC150 Dataset Summary Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example. We introduce a new dataset that includes queries that are out-of-scope (OOS), i.e., queries that do not fall into any of the system's supported intents. This poses a new challenge because models cannot assume that every query at inference… See the full description on the dataset page: https://huggingface.co/datasets/clinc/clinc_oos.texttext-classification10K<n<100K20 likes29k downloads3y agoHugging Face05clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes20k downloads4y agoHugging Face06LEAP /ClimSim_low-res-test1 likes19k downloads2y agoHugging Face07mteb /ClimateFEVER_test_top_250_only_w_correct-v2 ClimateFEVERHardNegatives An MTEB dataset Massive Text Embedding Benchmark CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Encyclopaedic, Written Reference https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.texttext-retrieval10K<n<100K0 likes16k downloads1y agoHugging Face08lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips 37,455 egocentric kitchen clips, one per narrated action, ready to explore in LightlyStudio. Search clips with natural language, browse them by narration, verb and noun, and spot clips whose narration doesn't match the video. 🚀 Explore it in LightlyStudio hf download lightly-ai/epic-kitchens-100-clips --repo-type dataset --local-dir epic-kitchens-100-clips cd epic-kitchens-100-clips pip install -r requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes16k downloads20h agoHugging Face09LEAP /ClimSim_high-resThe corresponding GitHub repo can be found here:https://github.com/leap-stc/ClimSim Read more: https://arxiv.org/abs/2306.08754. 13 likes16k downloads3y agoHugging Face10nvidia /Nemotron-ClimbLab ClimbLab Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.text-generation1B<n<10B38 likes14k downloads1y agoHugging Face11LEAP /ClimSim_low-res-expanded-test0 likes11k downloads2y agoHugging Face12clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes11k downloads4y agoHugging Face13clip-benchmark /wds_imagenet-aimage1K<n<10K0 likes9.7k downloads4y agoHugging Face14clip-benchmark /wds_imagenet1kimage10K<n<100K1 likes9k downloads4y agoHugging Face15BlueZeros /ClinicalAgentBenchMore detail about the dataset and the agentic framework can be found in https://github.com/BlueZeros/ReflecTool image100M<n<1B1 likes8.7k downloads1y agoHugging Face16LEAP /ClimSim_low-res-expandedThis is an expanded version of ClimSim_low-res. Each '.mlexpand.' file contains the same variables as in the corresponding '.mli.' file in ClimSim_low-res but also includes additional variables such as dynamical forcing, convection memory, cos/sin of latitude. Read more about these expanded input features at Section 6.3.3 in the SI of "ClimSim-Online: A Large Multi-scale Dataset and Framework for Hybrid ML-physics Climate Emulation": https://arxiv.org/abs/2306.08754. 0 likes8.5k downloads2y agoHugging Face17clip-benchmark /wds_imagenetv2image10K<n<100K0 likes7.7k downloads4y agoHugging Face18nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B130 likes7.7k downloads1y agoHugging Face19OptimalScale /ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.texttext-generation1B<n<10B16 likes7.3k downloads1y agoHugging Face20oumoumad /hdr-demo-clips HDR Demo Clips (Lightricks SDR→HDR) Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2). Each clip contains: hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred) sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred) thumbnail.jpg — 280px preview from the middle frame Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.imageimage-to-image10K<n<100K1 likes6.6k downloads4mo agoHugging Face21clips /beir-nl-cqadupstack Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.texttext-retrieval100K<n<1M0 likes6.2k downloads2y agoHugging Face22Kgshop /clientsimagen<1K0 likes5.1k downloads21h agoHugging Face23timbrooks /instructpix2pix-clip-filtered Dataset Card for InstructPix2Pix CLIP-filtered Dataset Summary The dataset can be used to train models to follow edit instructions. Edit instructions are available in the edit_prompt. original_image can be used with the edit_prompt and edited_image denotes the image after applying the edit_prompt on the original_image. Refer to the GitHub repository to know more about how this dataset can be used to train a model that can follow instructions. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.image100K<n<1M48 likes4.8k downloads4y agoHugging Face24FineEnvs /repo2rlenv-cli-gym Repo2RLEnv CLI-Gym Start with a healthy repository, synthesize a disruption and recovery, verify that damage breaks the original tests and recovery restores them, then export an environment-repair instruction and a deterministic verifier. Contains 25 Harbor tasks generated with the owned cli_gym recipe in Repo2RLEnv. Browse the complete task bundles in Harbor Visualiser or open the task folders. Each folder is a runnable Harbor task: tasks/<task_id>/ ├── task.toml… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-cli-gym.n<1K0 likes4.2k downloads12d agoHugging Face25gvlassis /ClimbMix ClimbMix About 🧗 A more convenient ClimbMix (https://arxiv.org/abs/2504.13161) Description Unfortunately, the original ClimbMix (https://huggingface.co/datasets/nvidia/ClimbMix) has four main inconveniences: It is in GPT2 tokens, meaning you have to detokenize it to inspect it or use it with another tokenizer. It contains all of the 20 clusters in order together (in the same "subset"), so you have to load the whole dataset in memory (~1TB) and shuffle it… See the full description on the dataset page: https://huggingface.co/datasets/gvlassis/ClimbMix.text100M<n<1B7 likes4.1k downloads1y agoHugging Face26clip-benchmark /wds_fer2013image10K<n<100K0 likes3.9k downloads4y agoHugging Face27EunsuKim /CLIcK CLIcK 🇰🇷🧠 A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean Introduction 🎉 CLIcK (Cultural and Linguistic Intelligence in Korean) is a comprehensive dataset designed to evaluate cultural and linguistic intelligence in the context of Korean language models. In an era where diverse language models are continually emerging, there is a pressing need for robust evaluation datasets, especially for non-English languages like Korean. CLIcK… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/CLIcK.textmultiple-choice1K<n<10K28 likes3.4k downloads2y agoHugging Face28CodedotAI /code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.text1M<n<10M20 likes3.2k downloads4y agoHugging Face29cambridge-climb /BabyLMDataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.10M<n<100M3 likes3k downloads2y agoHugging Face30RoboSynChallenge /cobotmagic_Sim_click_bellvideo1K<n<10K0 likes2.9k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.