datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maple
Overview
Maple is an open-source full-stack code dataset developed and released by Tudor Iustin.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.maplestory-worlds-creator-qa
MapleStory Worlds Creator QA
Synthetic question-answer dataset built from the official
MapleStory Worlds Creator Center
documentation. Questions are generated to be self-contained and grounded in the
source docs; answers avoid source/meta references so they read like an expert
explanation. Some QA pairs are composed from multiple related documents
(see combo_sources).
Parallel Korean/English. Intended for instruction tuning, QA, and retrieval.
Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.maple-personas
MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI
A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework.
Dataset Description
This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.maple-analyst-cap-sft-data
maple-analyst-cap-sft-data
Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B
ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato
TC (ThinkingCap).
Composición
Fuente
Filas
thinkingcap (curriculum, trazas bigbang)
1,782
openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0)
702
bigbang_mmlu
508
bigbang_bbh
441
hermes_function_calling
360
aya_dataset
342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Fabric AI.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.canada-china-trade
Canada-China B2B Trade Dataset
Dataset Description
A curated dataset of Canada-China bilateral trade statistics, commodity breakdowns, provincial data, and B2B sourcing knowledge for use in AI/LLM research and applications.
Maintained by: MapleBridge.io — AI-powered B2B matching platform for Canada-China trade.
Dataset Contents
File
Description
Rows
canada_china_trade_annual.csv
Annual bilateral trade volume 2015-2024 (CAD billions)
10… See the full description on the dataset page: https://huggingface.co/datasets/maplebridge/canada-china-trade.maplestory-worlds-creator-docs
MapleStory Worlds Creator Center Documentation
A curated dataset built from the official documentation of the
MapleStory Worlds Creator Center.
It is a parallel Korean/English documentation corpus intended for RAG, search,
embeddings, and domain language-model training.
The dataset covers all three Creator Center content types — guide documents
(doc), API Reference (api), and resources (res).
Composition
Document counts by type and language:
type
Description… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-docs.maplestory-worlds-creator-code-instruct
MapleStory Worlds Creator Code (mlua)
Instruction-style code dataset for mlua, the scripting language of
MapleStory Worlds. Built from the
official Creator Center example code: each example is grounded in its source
document and paired with a natural-language task, reasoning, a self-contained
explanation, and commented mlua code. Intended to teach LLMs to write mlua game
scripts.
The example code is preserved from the official source (a code-preservation check
rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.maple_nete_format_dataDataset for paper MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation
OracleGraphmedical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。maple_nete_format_dataDataset for MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation
Note that the train.index only contain the data index belonging to the train split (same for other splits); they need to work with reviews.pickle to obtain the data.
// since the file is very large, the compressed format is uploaded and huggingface cannot detect the splits
// adding splits here for clarity
- config_name: default
data_files:
- split: train
path:… See the full description on the dataset page: https://huggingface.co/datasets/NanaEilish/maple_nete_format_data.
