Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B123 likes637k downloads2y agoHugging Face02nyu-mll /glue Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.tabulartext-classification1M<n<10M1.1k likes519k downloads3y agoHugging Face03mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B320 likes482k downloads2y agoHugging Face04mlfoundations /dclm-pool-7b-2x3 likes231k downloads2y agoHugging Face05mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes97k downloads3y agoHugging Face06nyu-mll /blimp Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.texttext-classification10K<n<100K40 likes62k downloads3y agoHugging Face07mlfoundations /dcvlm_pool_large DCVLM-Pool (large) The raw candidate pool at the large scale of our DataComp-VLM benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as WebDataset tar shards — ≈4× the medium pool. 🚚 Upload in progress This repo is being populated incrementally and is not yet complete — shards are still being uploaded. Sources already present are final and safe to use; sources with fewer shards than the counts quoted below have not finished uploading yet.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_large.image-text-to-text1B<n<10B1 likes45k downloads2h agoHugging Face08mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes42k downloads3mo agoHugging Face09FineEnvs /HF_ML_Tasksmith HF ML Tasksmith Fifty PR-derived Harbor tasks from Accelerate, Diffusers, PEFT, Transformers and TRL, including CPU and GPU tasks. Contains 50 Harbor tasks generated with the owned tasksmith recipe in Repo2RLEnv. Browse the complete task bundles in Harbor Visualiser or open the task folders. Each folder is a runnable Harbor task: tasks/<task_id>/ ├── task.toml # Harbor configuration and provenance ├── instruction.md # Task shown to the coding agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith.n<1K4 likes42k downloads16d agoHugging Face10mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B57 likes39k downloads2y agoHugging Face11mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes37k downloads3y agoHugging Face12MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M286 likes35k downloads2y agoHugging Face13mlabonne /FineTome-100k FineTome-100k The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier. It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth". text100K<n<1M282 likes25k downloads2y agoHugging Face14mlabonne /harmful_behaviorstextn<1K164 likes25k downloads2y agoHugging Face15mueller91 /MLAADgated Introduction Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See the paper for more information. License MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted. Bibtex If you use this dataset, please consider citing it as follows. @article{muller2024mlaad, title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.audioaudio-classification100K<n<1M47 likes25k downloads1mo agoHugging Face16mlabonne /harmless_alpacatext10K<n<100K49 likes24k downloads2y agoHugging Face17mlfoundations /dclm-pool-1b-1x3 likes23k downloads2y agoHugging Face18mlfoundations /dcvlm_pool_small DCVLM-Pool (small) The raw candidate pool at the small scale of our DataComp-VLM benchmark: 120,940,134 samples / ~187.5B tokens / 10.3 TB across 166 source datasets, as WebDataset tar shards. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA DCVLM-baseline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small.image-text-to-text100M<n<1B0 likes22k downloads2mo agoHugging Face19mlfoundations /dcvlm_pool_medium DCVLM-Pool (medium) The raw candidate pool at the medium scale of our DataComp-VLM benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as WebDataset tar shards — ≈4× the small pool. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.image-text-to-text100M<n<1B0 likes22k downloads2mo agoHugging Face20lewtun /ml-intern-sessions ML Intern session traces This dataset contains ML Intern coding agent session traces uploaded from local ML Intern runs. The traces are stored as JSON Lines files under sessions/, with one file per session. Links ML Intern demo: https://smolagents-ml-intern.hf.space ML Intern CLI: https://github.com/huggingface/ml-intern Data description Each *.jsonl file contains a single ML Intern session converted to a Claude-Code-style event stream for the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/ml-intern-sessions.text-generation3 likes22k downloads1mo agoHugging Face21mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes21k downloads2mo agoHugging Face22nyu-mll /multi_nli Dataset Card for Multi-Genre Natural Language Inference (MultiNLI) Dataset Summary The Multi-Genre Natural Language Inference (MultiNLI) corpus is a crowd-sourced collection of 433k sentence pairs annotated with textual entailment information. The corpus is modeled on the SNLI corpus, but differs in that covers a range of genres of spoken and written text, and supports a distinctive cross-genre generalization evaluation. The corpus served as the basis for the shared task… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/multi_nli.texttext-classification100K<n<1M121 likes20k downloads3y agoHugging Face23mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes18k downloads3y agoHugging Face24mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face25mlfoundations /dclm-pool-7b-1x1 likes17k downloads2y agoHugging Face26MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes16k downloads2y agoHugging Face27jablonkagroup /chempile-mlift ChemPile-MLIFT A comprehensive multimodal dataset for chemistry property prediction using vision large language models 📋 Dataset Summary ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.imagetext-generation10M<n<100M14 likes15k downloads1y agoHugging Face28MLVU /MVLUgatedMLVU: Multi-task Long Video Understanding Benchmark This repo contains the annotation data and evaluation code for the paper "MLVU: A Comprehensive Benchmark for Multi-Task Long Video Understanding". 🔔 News: 🆕 7/28/2024: The data for the MLVU-Test set has been released (🤗 Link)! The test set includes 11 different tasks, featuring our newly added Sports Question Answering (SQA, single-detail LVU) and Tutorial… See the full description on the dataset page: https://huggingface.co/datasets/MLVU/MVLU.videoquestion-answering41 likes14k downloads2y agoHugging Face29japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes13k downloads2y agoHugging Face30VoiceOfML /MLMRL-Hub 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/discussions 提出。 此仓库存储马列毛主义与革命左翼仓储中心和图书馆的未重复资料(仅经一次md5检测,压缩包内容未去重。):https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/tree/main 。 请使用https://voiceofml-search.hf.space/MLMRL-Hub 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Hub )。 可使用:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub.audio10K<n<100K0 likes13k downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.