datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.Matrix
Matrix
An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models.
Dataset Composition
The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component:
Common Crawl
Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.COIG-CQIA
COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning
Dataset Details
Dataset Description
欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。
Welcome to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-CQIA.OProofs
OProofs
Formal Lean 4 theorem-proof pairs produced as part of the OProver project.
Fields
Field
Type
Description
formal_statement
string
Lean 4 theorem statement
formal_proof
string
Lean 4 proof body
cot_proof
string | null
Chain-of-thought reasoning preceding the proof, if available
prompt
string | null
Generation prompt, if available
Stats
Records: 6,804,694
Files: 73 parquet shards (zstd compressed)
Loading
from… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OProofs.MAPS
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. To balance performance and safety evaluation, our benchmark comprises 805 tasks: 405 from performance-oriented datasets (GAIA, SWE-bench, MATH) and 400 from the Agent Security… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS.MusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model
Dataset Card for MusicPile
MusicPile is the first pretraining corpus for developing musical abilities in large language models.
It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code.
You can easily load it:from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.AetherCode
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
Introduction
Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.MA-ProofBench
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis
English | 中文
We introduce MA-ProofBench, to the best of our knowledge, the first formal benchmark for evaluating large language models (LLMs) on theorem proving in Mathematical Analysis. It contains 200 rigorously formalized theorem-proving problems in Lean 4 + Mathlib (v4.28.0), split into two difficulty tiers:
Tier
Description
Source
Count
Level I
Undergraduate… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/MA-ProofBench.FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.TerminalTraj-5k-instances
TerminalTraj
🏆 ICML 2026 Spotlight
TerminalTraj is a dataset of high-quality, verified terminal agent trajectories generated from dockerized environments. It contains 50,733 verified terminal trajectories across eight domains.
GitHub Repository | Paper | Model (32B)
Dataset Description
Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/TerminalTraj-5k-instances.FineFineWeb-fasttext-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.FineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.MAPS_Verified
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Tudor Iustin.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.autonomous-driving-ethical-stability-accountability-mapping-v0.1
What this dataset tests
Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance.
Required outputs
stability impact description
accountability nodes
stability score
accountability score
recovery quality
Use case
Final layer of ethical navigation stack.
Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.m-a-p-FineFineWeb-sample
Unofficial m-a-p/FineFineWeb Sample
This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies.
Sampling Methodology
To create this subset, the following processing steps were taken:
Selection: 100 random .jsonl files were chosen from the original dataset.
Extraction: 10,000 rows were downloaded per selected file.
Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.maplestory-worlds-creator-qa
MapleStory Worlds Creator QA
Synthetic question-answer dataset built from the official
MapleStory Worlds Creator Center
documentation. Questions are generated to be self-contained and grounded in the
source docs; answers avoid source/meta references so they read like an expert
explanation. Some QA pairs are composed from multiple related documents
(see combo_sources).
Parallel Korean/English. Intended for instruction tuning, QA, and retrieval.
Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.maple-personas
MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI
A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework.
Dataset Description
This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.maple-analyst-cap-sft-data
maple-analyst-cap-sft-data
Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B
ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato
TC (ThinkingCap).
Composición
Fuente
Filas
thinkingcap (curriculum, trazas bigbang)
1,782
openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0)
702
bigbang_mmlu
508
bigbang_bbh
441
hermes_function_calling
360
aya_dataset
342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Fabric AI.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.deepshopper-mapper-reward-splits
DeepShopper frozen splits (mapper / reward)
Deterministic, leakage-safe train/test splits used across DeepShopper. Split assignment is a
stable sha1(need) hash (same need never crosses train/test; reproducible). Gender-stratified.
Contains: fashionrec_task1_mapper/{train,test}, amz_mapper/{female,male,other}.{train,test}
(need→plan), and amz_reward_bundle/{female,male,other}.{train,test} (need→outfit, reward
positives). Generated by scripts/make_mapper_reward_splits.py. Code:… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-mapper-reward-splits.map-territory-control-v01Cardinal Meta Dataset 3.3Map–Territory Control
Purpose
Test whether representations are not mistaken for reality
Test whether models, metrics, and frameworks are treated as tools
Test whether certainty is not imported from maps into territory
Central question
Is this a representation or the thing itself
What this dataset catches
Benchmark score treated as safety
Model prediction treated as outcome
Simulation treated as real world behavior
Framework treated as proof
Estimate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/map-territory-control-v01.hyperswitch-product-code-mapping
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 1801
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-product-code-mapping.canada-china-trade
Canada-China B2B Trade Dataset
Dataset Description
A curated dataset of Canada-China bilateral trade statistics, commodity breakdowns, provincial data, and B2B sourcing knowledge for use in AI/LLM research and applications.
Maintained by: MapleBridge.io — AI-powered B2B matching platform for Canada-China trade.
Dataset Contents
File
Description
Rows
canada_china_trade_annual.csv
Annual bilateral trade volume 2015-2024 (CAD billions)
10… See the full description on the dataset page: https://huggingface.co/datasets/maplebridge/canada-china-trade.maplestory-worlds-creator-docs
MapleStory Worlds Creator Center Documentation
A curated dataset built from the official documentation of the
MapleStory Worlds Creator Center.
It is a parallel Korean/English documentation corpus intended for RAG, search,
embeddings, and domain language-model training.
The dataset covers all three Creator Center content types — guide documents
(doc), API Reference (api), and resources (res).
Composition
Document counts by type and language:
type
Description… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-docs.MAPF-FrozenLake-Benchmark
MAPF-FrozenLake Benchmark
Evaluation benchmark for the paper
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
Three configs (benchmark_wr025 / benchmark_wr050 / benchmark_wr075)
correspond to the wait-ratio threshold of the underlying CBS-optimal
solution (higher = more inter-agent coordination required).
Each config has three splits by agent count.
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/MAPF-FrozenLake-Benchmark.maplestory-worlds-creator-code-instruct
MapleStory Worlds Creator Code (mlua)
Instruction-style code dataset for mlua, the scripting language of
MapleStory Worlds. Built from the
official Creator Center example code: each example is grounded in its source
document and paired with a natural-language task, reasoning, a self-contained
explanation, and commented mlua code. Intended to teach LLMs to write mlua game
scripts.
The example code is preserved from the official source (a code-preservation check
rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.clinical-quad-recruitment-coherence-mapping-suite-v0.1Clarus Clinical Quad Coupling Recruitment Coherence Mapping Suite v0.1
What this dataset isThis dataset tests whether a model can detect recruitment incoherence under four-node coupling pressure.
Quad coupling nodes
Biological eligibility definition
Concomitant medication or background therapy filters
Operational measurement and site process variance
Governance constraints limiting protocol flexibility
Input
One recruitment vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-coherence-mapping-suite-v0.1.
