Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B123 likes637k downloads2y agoHugging Face02nyu-mll /glue Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.tabulartext-classification1M<n<10M1.1k likes519k downloads3y agoHugging Face03mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B320 likes482k downloads2y agoHugging Face04nyu-mll /blimp Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.texttext-classification10K<n<100K40 likes62k downloads3y agoHugging Face05mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes42k downloads3mo agoHugging Face06mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B57 likes39k downloads2y agoHugging Face07mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes37k downloads3y agoHugging Face08MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M286 likes35k downloads2y agoHugging Face09mlabonne /FineTome-100k FineTome-100k The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier. It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth". text100K<n<1M282 likes25k downloads2y agoHugging Face10mlabonne /harmful_behaviorstextn<1K164 likes25k downloads2y agoHugging Face11mlabonne /harmless_alpacatext10K<n<100K49 likes24k downloads2y agoHugging Face12mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes21k downloads2mo agoHugging Face13nyu-mll /multi_nli Dataset Card for Multi-Genre Natural Language Inference (MultiNLI) Dataset Summary The Multi-Genre Natural Language Inference (MultiNLI) corpus is a crowd-sourced collection of 433k sentence pairs annotated with textual entailment information. The corpus is modeled on the SNLI corpus, but differs in that covers a range of genres of spoken and written text, and supports a distinctive cross-genre generalization evaluation. The corpus served as the basis for the shared task… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/multi_nli.texttext-classification100K<n<1M121 likes20k downloads3y agoHugging Face14mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes18k downloads3y agoHugging Face15mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face16jablonkagroup /chempile-mlift ChemPile-MLIFT A comprehensive multimodal dataset for chemistry property prediction using vision large language models 📋 Dataset Summary ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.imagetext-generation10M<n<100M14 likes15k downloads1y agoHugging Face17LeMaterial /LeMat-Bulk-MLIP-Hull LeMat-Bulk MLIP Hull Reference Datasets This dataset contains materials close to the convex hull computed using various ML interatomic potentials (MLIPs). Dataset Splits all: Contains ALL materials with hull energies for all MLIPs (no threshold filtering) dft, orb, uma, mace_mp, mace_omat: Materials within 0.001 eV/atom of respective hulls Energy Types dft: DFT reference energies orb: ORB model energies uma: UMA model energies mace_mp: MACE-MP model energies… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull.tabular1M<n<10M0 likes11k downloads1y agoHugging Face18mlfoundations-dev /aops_forum_filteredtext10K<n<100K0 likes10k downloads2y agoHugging Face19inductionlabs /frontier-ml-tasks Frontier MLE tasks Data for the tasks in induction-labs/frontier-ml. Each task has one folder: <slug>/public/ given to the coding agent verbatim, at /task/public <slug>/private/ held-out data for the trusted verifier, at /tests/private <slug>/artifacts/ optional starting files for the agent, at /artifacts <slug>/manifest.json optional provenance summary Tasks pin this repository by commit in tasks/<slug>/task.yaml. Starting model weights are downloaded from their… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/frontier-ml-tasks.imagen<1K0 likes8.8k downloads17h agoHugging Face20mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes8.7k downloads2y agoHugging Face21okashurin /fineweb_10bt_ml_64text100M<n<1B0 likes8.3k downloads7mo agoHugging Face22modelomics /gh-ml GitHub ML A continually refreshed registry of GitHub repositories across ML fields. The curated current view applies ml-contribution-v5 and seeks projects that present a distinct contribution to an ML model, method, or technique. The broader candidates view includes current-view rows and repositories with heuristic evidence of probable original ML content, as well as established qualified review cases. Candidate eligibility is independent of strict selection status: a row… See the full description on the dataset page: https://huggingface.co/datasets/modelomics/gh-ml.tabular1M<n<10M0 likes8.1k downloads45m agoHugging Face23okashurin /fineweb_10bt_ml_16text100M<n<1B0 likes7.8k downloads7mo agoHugging Face24sy1998 /MLVU_devtext1K<n<10K2 likes7.4k downloads2y agoHugging Face25mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes6.7k downloads2y agoHugging Face26VoiceOfML /MLMRL-Library 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/discussions 提出。 此仓库存储重要书库备份:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/tree/main 。 请使用:https://voiceofml-search.hf.space/MLMRL-Library 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Library )。 可使用:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Library.document1 likes6.5k downloads7d agoHugging Face27GD-ML /MAPBench-V2For more details, please check our project page. Paper: https://arxiv.org/abs/2601.05432 Repository: https://github.com/AMAP-ML/Thinking-with-Map image1K<n<10K4 likes6.5k downloads8mo agoHugging Face28mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes6.4k downloads2y agoHugging Face29mlx-community /mlx-model-explorer-data MLX Model Explorer Data An anonymous record of how people use MLX Model Explorer to choose an MLX model for their Mac: which model families, sizes, quantizations, memory classes and context lengths they look at, and which models they go on to open, compare or download. It also holds the community reports ("it worked", "too slow") and real MLX benchmark results that people choose to contribute. The goal is to answer, with data: what is the MLX community actually trying to run… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/mlx-model-explorer-data.tabular1K<n<10K4 likes5.7k downloads32m agoHugging Face30mlfoundations-dev /Eurus-2-7B-SFT_eval_2e29 mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 2.3 21.0 30.6 11.0 11.4 10.4 6.8 1.5 2.1 1.3 4.1 4.4 AIME24 Average Accuracy: 2.33% ± 0.67% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 0.00% 0 30 2 3.33% 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.tabular1K<n<10K0 likes5.6k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.