Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.tabulartext-classification1B<n<10B205 likes4.3m downloads2y agoHugging Face02m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face03m-a-p /Matrix Matrix An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models. Dataset Composition The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component: Common Crawl Extracts from the Common Crawl project, featuring a rich diversity of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.texttext-generation1B<n<10B176 likes33k downloads2y agoHugging Face04m-a-p /COIG-CQIA COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning Dataset Details Dataset Description 欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。 Welcome to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/COIG-CQIA.textquestion-answering10K<n<100K777 likes7.9k downloads2y agoHugging Face05m-a-p /OProofs OProofs Formal Lean 4 theorem-proof pairs produced as part of the OProver project. Fields Field Type Description formal_statement string Lean 4 theorem statement formal_proof string Lean 4 proof body cot_proof string | null Chain-of-thought reasoning preceding the proof, if available prompt string | null Generation prompt, if available Stats Records: 6,804,694 Files: 73 parquet shards (zstd compressed) Loading from… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OProofs.texttext-generation1M<n<10M4 likes5.8k downloads5mo agoHugging Face06Fujitsu-FRE /MAPS Dataset Card for Multilingual Benchmark for Global Agent Performance and Security This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. To balance performance and safety evaluation, our benchmark comprises 805 tasks: 405 from performance-oriented datasets (GAIA, SWE-bench, MATH) and 400 from the Agent Security… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS.texttext-generation1K<n<10K13 likes1k downloads6mo agoHugging Face07m-a-p /MusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model Dataset Card for MusicPile MusicPile is the first pretraining corpus for developing musical abilities in large language models. It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code. You can easily load it:from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.texttext-generation1M<n<10M60 likes996 downloads2y agoHugging Face08m-a-p /AetherCode AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions Introduction Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.tabulartext-generationn<1K8 likes576 downloads1y agoHugging Face09openbmb /MA-ProofBench MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis English | 中文 We introduce MA-ProofBench, to the best of our knowledge, the first formal benchmark for evaluating large language models (LLMs) on theorem proving in Mathematical Analysis. It contains 200 rigorously formalized theorem-proving problems in Lean 4 + Mathlib (v4.28.0), split into two difficulty tiers: Tier Description Source Count Level I Undergraduate… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/MA-ProofBench.texttext-generationn<1K16 likes505 downloads2mo agoHugging Face10m-a-p /FineFineWeb-test FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.tabulartext-classification1M<n<10M5 likes501 downloads2y agoHugging Face11m-a-p /TerminalTraj-5k-instances TerminalTraj 🏆 ICML 2026 Spotlight TerminalTraj is a dataset of high-quality, verified terminal agent trajectories generated from dockerized environments. It contains 50,733 verified terminal trajectories across eight domains. GitHub Repository | Paper | Model (32B) Dataset Description Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/TerminalTraj-5k-instances.text-generation6 likes470 downloads2mo agoHugging Face12m-a-p /FineFineWeb-fasttext-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-fasttext-seeddata.text-classificationn>1T0 likes415 downloads2y agoHugging Face13m-a-p /FineFineWeb-validation FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.tabulartext-classification10K<n<100K1 likes367 downloads2y agoHugging Face14m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes332 downloads2y agoHugging Face15Fujitsu-FRE /MAPS_Verified Dataset Card for Multilingual Benchmark for Global Agent Performance and Security This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.texttext-generation1K<n<10K3 likes215 downloads8mo agoHugging Face16tudor-iustin22 /maple Overview Maple is an open-source full-stack code dataset developed and released by Tudor Iustin. It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems. Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.texttext-generation10K<n<100K0 likes128 downloads3mo agoHugging Face17ClarusC64 /autonomous-driving-ethical-stability-accountability-mapping-v0.1 What this dataset tests Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance. Required outputs stability impact description accountability nodes stability score accountability score recovery quality Use case Final layer of ethical navigation stack. Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.tabulartext-generationn<1K0 likes106 downloads8mo agoHugging Face18agentlans /m-a-p-FineFineWeb-sample Unofficial m-a-p/FineFineWeb Sample This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies. Sampling Methodology To create this subset, the following processing steps were taken: Selection: 100 random .jsonl files were chosen from the original dataset. Extraction: 10,000 rows were downloaded per selected file. Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.tabulartext-classification1M<n<10M0 likes96 downloads4mo agoHugging Face19msw-ai-tf /maplestory-worlds-creator-qa MapleStory Worlds Creator QA Synthetic question-answer dataset built from the official MapleStory Worlds Creator Center documentation. Questions are generated to be self-contained and grounded in the source docs; answers avoid source/meta references so they read like an expert explanation. Some QA pairs are composed from multiple related documents (see combo_sources). Parallel Korean/English. Intended for instruction tuning, QA, and retrieval. Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.textquestion-answering100K<n<1M2 likes46 downloads3mo agoHugging Face20prdeepakbabu /maple-personas MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework. Dataset Description This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.texttext-generation1K<n<10K0 likes45 downloads8mo agoHugging Face21Davd-b01 /maple-analyst-cap-sft-data maple-analyst-cap-sft-data Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato TC (ThinkingCap). Composición Fuente Filas thinkingcap (curriculum, trazas bigbang) 1,782 openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0) 702 bigbang_mmlu 508 bigbang_bbh 441 hermes_function_calling 360 aya_dataset 342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.texttext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face22FabricAI /maple Overview Maple is an open-source full-stack code dataset developed and released by Fabric AI. It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems. Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.text-generation10K<n<100K0 likes33 downloads2mo agoHugging Face23flavianv /deepshopper-mapper-reward-splits DeepShopper frozen splits (mapper / reward) Deterministic, leakage-safe train/test splits used across DeepShopper. Split assignment is a stable sha1(need) hash (same need never crosses train/test; reproducible). Gender-stratified. Contains: fashionrec_task1_mapper/{train,test}, amz_mapper/{female,male,other}.{train,test} (need→plan), and amz_reward_bundle/{female,male,other}.{train,test} (need→outfit, reward positives). Generated by scripts/make_mapper_reward_splits.py. Code:… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-mapper-reward-splits.text-generation0 likes31 downloads3mo agoHugging Face24ClarusC64 /map-territory-control-v01Cardinal Meta Dataset 3.3Map–Territory Control Purpose Test whether representations are not mistaken for reality Test whether models, metrics, and frameworks are treated as tools Test whether certainty is not imported from maps into territory Central question Is this a representation or the thing itself What this dataset catches Benchmark score treated as safety Model prediction treated as outcome Simulation treated as real world behavior Framework treated as proof Estimate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/map-territory-control-v01.texttext-generationn<1K0 likes29 downloads9mo agoHugging Face25archit11 /hyperswitch-product-code-mapping Rust Commit Dataset - Hyperswitch Dataset Description This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository. Dataset Summary Total Examples: 1801 Language: Rust Source: Hyperswitch GitHub repository Format: Prompt-response pairs for supervised fine-tuning (SFT) Data Fields prompt: The commit message describing the change response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-product-code-mapping.texttext-generation1K<n<10K0 likes28 downloads1y agoHugging Face26maplebridge /canada-china-trade Canada-China B2B Trade Dataset Dataset Description A curated dataset of Canada-China bilateral trade statistics, commodity breakdowns, provincial data, and B2B sourcing knowledge for use in AI/LLM research and applications. Maintained by: MapleBridge.io — AI-powered B2B matching platform for Canada-China trade. Dataset Contents File Description Rows canada_china_trade_annual.csv Annual bilateral trade volume 2015-2024 (CAD billions) 10… See the full description on the dataset page: https://huggingface.co/datasets/maplebridge/canada-china-trade.text-generationn<1K0 likes28 downloads7mo agoHugging Face27msw-ai-tf /maplestory-worlds-creator-docs MapleStory Worlds Creator Center Documentation A curated dataset built from the official documentation of the MapleStory Worlds Creator Center. It is a parallel Korean/English documentation corpus intended for RAG, search, embeddings, and domain language-model training. The dataset covers all three Creator Center content types — guide documents (doc), API Reference (api), and resources (res). Composition Document counts by type and language: type Description… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-docs.texttext-generation1K<n<10K2 likes26 downloads4mo agoHugging Face28LARK-Lab /MAPF-FrozenLake-Benchmark MAPF-FrozenLake Benchmark Evaluation benchmark for the paper From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Three configs (benchmark_wr025 / benchmark_wr050 / benchmark_wr075) correspond to the wait-ratio threshold of the underlying CBS-optimal solution (higher = more inter-agent coordination required). Each config has three splits by agent count. Load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/MAPF-FrozenLake-Benchmark.texttext-generation1K<n<10K1 likes24 downloads4mo agoHugging Face29msw-ai-tf /maplestory-worlds-creator-code-instruct MapleStory Worlds Creator Code (mlua) Instruction-style code dataset for mlua, the scripting language of MapleStory Worlds. Built from the official Creator Center example code: each example is grounded in its source document and paired with a natural-language task, reasoning, a self-contained explanation, and commented mlua code. Intended to teach LLMs to write mlua game scripts. The example code is preserved from the official source (a code-preservation check rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.texttext-generation1K<n<10K2 likes24 downloads3mo agoHugging Face30ClarusC64 /clinical-quad-recruitment-coherence-mapping-suite-v0.1Clarus Clinical Quad Coupling Recruitment Coherence Mapping Suite v0.1 What this dataset isThis dataset tests whether a model can detect recruitment incoherence under four-node coupling pressure. Quad coupling nodes Biological eligibility definition Concomitant medication or background therapy filters Operational measurement and site process variance Governance constraints limiting protocol flexibility Input One recruitment vignette OutputReturn strict JSON only. Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-coherence-mapping-suite-v0.1.texttext-generationn<1K0 likes22 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.