datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moe-demo-clean
RoboTwin MOE Demo Clean
Raw RoboTwin demonstration data copied from bos:/lab-test/moe-demo-clean/.
The dataset is organized by task directories such as
place_can_basket-demo_clean-200/. Each task directory contains episode
subdirectories with an HDF5 trajectory file and an instructions.json file.
story_clozeTinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.moeMoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.moe-routing-drift-results
MoE routing drift — results
Measurements for a 2x2 experiment: adaptation (none / GEPA / prompt-tuning /
prefix-tuning) crossed with router retraining (frozen gate / gate retrained), on
inclusionAI/Ling-mini-2.0 and Qwen/Qwen3-30B-A3B-Instruct-2507. The weights are in
moe-routing-drift-checkpoints.
Content warning. quality/*/*.responses.jsonl contain verbatim comments from
civil_comments together with model outputs; the task is toxicity labelling, so the text
includes insults… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/moe-routing-drift-results.VideoVista-CulturalLingo
VideoVista-CulturalLingo
This repository contains the VideoVista-CulturalLingo, introduced in VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages,
and Domains in Video Comprehension.
🎉 Our new VideoVista-CulturalLingo bridges cultures (China, North America, and Europe), languages (Chinese and English), and domains (140+)in video comprehension.
🌍 Welcome to join us on this journey of video understanding!
🔥 News
[2025/11/17]… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista-CulturalLingo.glm-moe-dsa-tiny-cpu-repro-v1
Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay
Reproducibility evidence for
malaiwah/glm-moe-dsa-tiny-random-bf16,
checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37.
This is a synthetic pipeline test, not a quality benchmark, quantization measurement,
qualified production reference, or registry submission. The model is random-init.
No GPU or paid cloud job was used.
Observed result
Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.Muice-Dataset
Muice-Dataset
沐雪角色扮演训练集
🤖ModelScope|
🤗HuggingFace|
(Github)Muicebot
更新日志
2026.05.18: 因为作者的论文使用到了本训练集需要引用,故更新 DOI 引用
2026.02.05: 小型更新,此次更新过后不再有新的数据集产生。
2025.08.23: 完整开源所有训练集以作研究用途,大幅更新自述文件
2025.02.14: 更新测试集以便透明化测试流程
2025.01.29: 新年快乐!为了感谢大家对沐雪训练集的喜欢,我们重写了训练集并额外提供 500 条训练集给大家。你可以在 这里 查看训练集重写目的和具体内容。除此之外,我们用 Sharegpt 格式规范了训练集格式,现在应该不会那么容易报错了...我们期望大家合理使用我们的训练集并训练出更高质量的模型,祝各位生活愉快。
简介… See the full description on the dataset page: https://huggingface.co/datasets/Moemu/Muice-Dataset.all_datasets_Qwen3-30B-A3B_moe_patternsMoeGirlPedia_zh_cleaned_latest
🌐Language 中文|English
本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。
特色
⚡体积优势
🧠文本易理解
💬更符合中文语境
仅经过基础清洗的数据集
1.06GB
存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚
暴力抹除非中文文字,导致信息缺失严重
本数据集
0.74GB(30.2%↓)
通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语)
仅经过基础清洗的数据集
size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。
'''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}}
url=akanejunpei.jpg|position=up
图片说明=999中的茜(2027,21岁)
|本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.Vision-Flan-191-1kKeural-MoE-14B-stage1-DatasetPredicting-Experts-for-MOE-Suite
Predicting Experts for MoE: A Coverage-First Routing Benchmark
Goal: predict the complete set of experts that a future
Mixture-of-Experts layer will activate, using only information that is
causally available before that layer executes.
This dataset turns expert prefetch prediction into a standalone machine-learning
problem. It contains 98,292 routed generated tokens and 7,371,900 ordered
expert-route labels from 12 synthetic, realistic coding tasks evaluated with… See the full description on the dataset page: https://huggingface.co/datasets/mistrjirka/Predicting-Experts-for-MOE-Suite.uniagent-qwen3-30b-a3b-r2e-rollouts-r2e_moe_maxrl_09111923MSVQAUnofficial training-ready fork of Kaij00/MSVQA.
wmt16_Qwen3-30B-A3B_moe_patternswindows-rtx-4060ti-8gb-moe-offload-bench-2026-05
RTX 4060 Ti 8GB — Multi-Model Benchmark (2026-05)
practitioner benchmarks on consumer hardware (8GB VRAM, 32GB RAM). 10 models tested, covering MoE expert offload, hybrid SSM architectures, dense models, MLA, dense partial GPU offload, and the 1B speed ceiling. all runs on the same physical rig, same methodology.
current leaderboard (decode tok/s at sweet spot)
model
active params
GGUF size
sweet spot tok/s
quality (6 tests)
architecture
Llama 3.2 1B
1.24B
771… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05.xsum_deepseek-moe-16b-chat_token_patternsall_datasets_deepseek-moe-16b-chat_moe_patternsllm-selection-dataset
llm-selection-dataset
Selected training samples in Parquet format.
Each file contains dataset, id, and messages. Each message contains role and content.
The align_<dimension>_<threshold>.parquet files contain selections for the corresponding alignment dimension and minimum score threshold.
gpqa_diamond_Qwen3-30B-A3B_moe_patternsless-is-moe-s1-calibration-128-seq8192
Less-is-MoE S1K calibration data — 128 samples, seq_length 8192
This is the fixed calibration artifact used to prune GPT-OSS-120B,
Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant:
yentinglin/s1K-1.1-trl-format revision
58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by
Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows.
For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.aime2024_Qwen3-30B-A3B_moe_patternswikihopverlmath-500_Qwen3-30B-A3B_moe_patternsaime2024_Qwen3-30B-A3B_moe_patterns_logitsmoe-unified-dataset-sota
moe-unified-dataset-sota
A unified dataset for training Mixture of Experts (MoE) models, combining multiple high-quality sources.
Dataset Statistics
Total Examples: 2,186,763
Train Split: 2,077,424
Test Split: 109,339
Sources
NousResearch/Hermes-3-Dataset - General instruction following, math, coding (~950k examples)
Salesforce/xlam-function-calling-60k - Function/tool calling (60k examples)
MegaScience/TextbookReasoning - Academic Q&A (~650k examples)… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/moe-unified-dataset-sota.
