datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.PIN-200M
PIN-200M
A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents"
Paper: https://arxiv.org/abs/2406.13923
This dataset contains around 200M samples in PIN format, with around 312 TB storage.
🚀 News
[ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it.
[ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-200M.FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.Matrix
Matrix
An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models.
Dataset Composition
The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component:
Common Crawl
Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
OpenCodeInterpreter
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.PIN-14M
PIN-14M
A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents"
Paper: https://arxiv.org/abs/2406.13923
This dataset contains 14M samples in PIN format, with around 18.79 TB storage.
🚀 News
[ 2025.09.04 ] !NEW! 🔥 We have completed the final version of the PIN-14M dataset and conducted some simple statistics on it.
[ 2024.12.12 ] !NEW! 🔥 We have updated the quality signals for all subsets, with the dataset now containing 7.33B tokens… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-14M.SuperGPQAThis repository contains the data presented in SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.
Tutorials for submitting to the official leadboard
coming soon
📜 License
SuperGPQA is a composite dataset that includes both original content and portions of data derived from other sources. The dataset is made available under the Open Data Commons Attribution License (ODC-BY), which asserts no copyright over the underlying content.
This means that while the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SuperGPQA.COIG-CQIA
COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning
Dataset Details
Dataset Description
欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。
Welcome to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-CQIA.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
OProofs
OProofs
Formal Lean 4 theorem-proof pairs produced as part of the OProver project.
Fields
Field
Type
Description
formal_statement
string
Lean 4 theorem statement
formal_proof
string
Lean 4 proof body
cot_proof
string | null
Chain-of-thought reasoning preceding the proof, if available
prompt
string | null
Generation prompt, if available
Stats
Records: 6,804,694
Files: 73 parquet shards (zstd compressed)
Loading
from… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OProofs.MapPool
MapPool - Bubbling up an extremely large corpus of maps for AI
MapPool is a dataset of 75 million potential maps and textual captions. It has been derived from CommonPool, a dataset consisting of 12 billion text-image pairs from the Internet. The images have been encoded by a vision transformer and classified into maps and non-maps by a support vector machine. This approach outperforms previous models and yields a validation accuracy of 98.5%. The MapPool dataset may help to train… See the full description on the dataset page: https://huggingface.co/datasets/sraimund/MapPool.42_map_datasetMAP-CC
MAP-CC
🌐 Homepage | 🤗 MAP-CC | 🤗 CHC-Bench | 🤗 CT-LLM | 📖 arXiv | GitHub
An open-source Chinese pretraining dataset with a scale of 800 billion tokens, offering the NLP community high-quality Chinese pretraining data.
Disclaimer
This model, developed for academic purposes, employs rigorously compliance-checked training data to uphold the highest standards of integrity and compliance. Despite our efforts, the inherent complexities of data and the broad spectrum of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MAP-CC.SciMMIR
Dataset Card for "SciMMIR_dataset"
SciMMIR
This is the repo for the paper SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval.
In this paper, we propose a novel SciMMIR benchmark and a corresponding dataset designed to address the gap in evaluating multi-modal information retrieval (MMIR) models in the scientific domain.
It is worth mentioning that we define a data hierarchical architecture of "Two subsets, Five subcategories" and use human-created… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SciMMIR.WildSongBench🤗 WildSongBench
A benchmark for full-song music generation
192 prompts · 94 Chinese · 98 English
🎵 YuE2 project
·
🚀 Quick start
·
📊 Benchmarks
·
🔁 Reproduce
·
📄 arXiv
·
PDF
·
📚 Citation
WildSongBench (WSB) contains 192 song-generation prompts: 94 Chinese and 98 English, used in the YuE2 benchmarks. This repository provides prompts, exact… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/WildSongBench.cleanvid-15m_map
CleanVid Map (15M) 🎥
TempoFunk Video Generation Project
CleanVid-15M is a large-scale dataset of videos with multiple metadata entries such as:
Textual Descriptions 📃
Recording Equipment 📹
Categories 🔠
Framerate 🎞️
Aspect Ratio 📺
CleanVid aim is to improve the quality of WebVid-10M dataset by adding more data and cleaning the dataset by dewatermarking the videos in it.
This dataset includes only the map with the urls and metadata, with 3,694,510 more entries than… See the full description on the dataset page: https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map.war3-maps
Warcraft III Community Map Archive
This public dataset preserves community-created Warcraft III maps and campaigns
for interoperability testing, search, research, and long-term access. Files are
deduplicated by SHA-256. Titles and other metadata are extracted with
war3-manager where the format permits.
Search and download individual maps: https://war3-archive.github.io/war3-maps/
Source and issue tracker: https://github.com/war3-archive/war3-maps
Layout… See the full description on the dataset page: https://huggingface.co/datasets/magicwenli/war3-maps.llmtcl
⚡ LitGPT
20+ high-performance LLMs with recipes to pretrain, finetune, and deploy at scale.
✅ From scratch implementations ✅ No abstractions ✅ Beginner friendly
✅ Flash attention ✅ FSDP ✅ LoRA, QLoRA, Adapter
✅ Reduce GPU memory (fp4/8/16/32) ✅ 1-1000+ GPUs/TPUs ✅ 20+ LLMs
Quick start •
Models •
Finetune •
Deploy •
All workflows •
Features •
Recipes (YAML) •
Lightning AI •
Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/Maple222/llmtcl.MapTrace
MapTrace: A 2M-Sample Synthetic Dataset for Path Tracing on Maps
Welcome to the MapTrace dataset! If you use this dataset in your work, please cite our paper below.
For more details about our methodology and findings, please visit our project page or read the official white paper.
This work was also recently featured on the Google Research Blog.
Code & Scripts
Official training and data loading scripts are available in our GitHub repository:… See the full description on the dataset page: https://huggingface.co/datasets/google/MapTrace.tvelve_map_complete_datasetCode-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
Introduction
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and related… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Code-Feedback.SimpleVQA
SimpleVQA
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
Dataset: https://huggingface.co/datasets/m-a-p/SimpleVQA
Abstract
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work, we… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SimpleVQA.model-mappingADFA_MappingCMMMU
CMMMU
🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | GitHub
Introduction
CMMMU includes 12k manually collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering, like its companion, MMMU. These questions span 30 subjects and comprise 39 highly heterogeneous image types, such as charts, diagrams, maps, tables, music… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CMMMU.DocCornerDataset
DocCornerDataset
A high-quality document corner detection dataset for training models to detect the four corners of documents in images. This dataset is optimized for building robust document scanning and perspective correction applications.
Dataset Examples
Training Set
Validation Set
Test Set
Green polygons show the annotated document corners
Dataset Description
This dataset contains images with document corner… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/DocCornerDataset.MAPS
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. To balance performance and safety evaluation, our benchmark comprises 805 tasks: 405 from performance-oriented datasets (GAIA, SWE-bench, MATH) and 400 from the Agent Security… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS.OmniBench
OmniBench
🌐 Homepage | 🏆 Leaderboard | 📖 Arxiv Paper | 🤗 Paper | 🤗 OmniBench Dataset | | 🤗 OmniInstruct_V1 Dataset | 🦜 Tweets
The project introduces OmniBench, a novel benchmark designed to rigorously evaluate models' ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define models capable of such tri-modal processing as omni-language models (OLMs).
Mini Leaderboard
This table shows the omni-language models in… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OmniBench.MusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model
Dataset Card for MusicPile
MusicPile is the first pretraining corpus for developing musical abilities in large language models.
It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code.
You can easily load it:from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.dalton-lv-meshes
Dalton et al. (2023) Left-Ventricle Meshes — DOLFINx XDMF
Finite-element meshes of 2967 synthetic left-ventricle (LV) geometries, converted
from the raw node data of
D. Dalton, H. Gao, D. Husmeier, "Emulation of cardiac mechanics using Graph
Neural Networks", Computer Methods in Applied Mechanics and Engineering 401
(2023) 115645. Dataset: Zenodo 7075055
· Code: https://github.com/dodaltuin/passive-lv-gnn-emul
The upstream download contains only a node-adjacency graph plus… See the full description on the dataset page: https://huggingface.co/datasets/mapengfei/dalton-lv-meshes.
