datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
AndroidCodeSuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.optiq-code-traces
OptiQ Code Traces
Gold-verified agentic software-engineering trajectories, produced by OptiQ Code, the terminal coding agent for local models on a Mac. Each trajectory is a full tool-calling run against a real repository bug, and every resolved label is set by executing the gold tests (FAIL_TO_PASS + PASS_TO_PASS) after applying the model's patch, never by the agent's own self-report.
The dataset is 1,789 agent sessions in HuggingFace Session-Traces format (the agent-traces… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-code-traces.Logics-SWE-Env-2.5K
Logics-SWE-Env-2.5K
2,553 software engineering task instances · 1,771 repositories · 4 programming languages
🤗 Related model: Logics-SWE-Qwen3.6-27B
📄 Paper: One to More, More to One
💻 GitHub: AgenticBigBang
Overview
What is this dataset?
Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.ml-intern-sessions
ML Intern session traces
This dataset contains ML Intern coding agent session traces uploaded from local
ML Intern runs. The traces are stored as JSON Lines files under sessions/,
with one file per session.
Links
ML Intern demo: https://smolagents-ml-intern.hf.space
ML Intern CLI: https://github.com/huggingface/ml-intern
Data description
Each *.jsonl file contains a single ML Intern session converted to a
Claude-Code-style event stream for the… See the full description on the dataset page: https://huggingface.co/datasets/clem/ml-intern-sessions.MAPBench-V1For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
MLLM_pathMLV-Bench
MedHorizon / MLV-Bench
MedHorizon, also released as MLV-Bench, is a long-context medical video benchmark for evaluating multimodal models on full-procedure clinical videos. The benchmark emphasizes two properties that are not captured by short-clip medical video datasets: extremely sparse evidence retrieval and multi-hop reasoning over observations distributed across a full procedure.
Dataset Contents
Videos: 340 full-procedure videos.
Questions: 1,253 multiple-choice QA… See the full description on the dataset page: https://huggingface.co/datasets/DBD123/MLV-Bench.jsonl-mls-hubert_large_ll60k-layer_22MLVU-Attention-Review
MLVU Attention Review 下载与解压说明
本仓库提供一个未压缩 TAR,里面是完整静态 HTML 展示、所有页面所需图片、QA/GT、模型原始回答、attention 元数据与未归一化 grids.npz,并附离线查看和逐文件 SHA256 校验脚本。归档不含模型权重、完整源视频或推理环境。
容量要求
下载和解压期间需同时存放 TAR 与解压内容,请预留至少归档大小约 2.2 倍的可用空间。
文件系统必须支持大于 4 GB 的单文件(NTFS、exFAT、ext4、APFS 等;FAT32 不可用)。
Windows 建议在较短路径中操作,例如 D:\reviews\MLVU,避免长路径限制。
下载
安装最新版 Hugging Face CLI:
pip install -U huggingface_hub
hf download GraciaChen/MLVU-Attention-Review MLVU-Attention-Review.tar SHA256SUMS… See the full description on the dataset page: https://huggingface.co/datasets/GraciaChen/MLVU-Attention-Review.seeclick-web-commercial-mlx
SeeClick Web Commercial Dataset (MLX-VLM Format)
Commercial-use friendly GUI grounding dataset from SeeClick Web data.
Apache 2.0 licensed - safe for commercial applications.
Dataset Description
This dataset contains ~20k examples for training Vision-Language Models to predict
click coordinates given a screenshot and instruction. Derived from SeeClick Web
crawled data (Apache 2.0).
Key Features
License: Apache 2.0 (commercial use allowed)
Format: MLX-VLM… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/seeclick-web-commercial-mlx.fineweb-2-ml4mmle-trajectories
MLE-lite trajectories: sol and Opus5
This run contains 39 Harbor ATIF-v1.7 trajectories from a 21-task MLE-lite evaluation of two native coding agents, plus 42 outcome records and official grading reports. Tasks ran on 2026-09-19–20 using Harbor 0.23.0 with a direct Kubernetes backend.
Model
Native agent
Tasks
Valid submissions / trajectories
Gold / silver / bronze
Any medal
gpt-5.6-sol
Codex 0.155.1
21
21
11 / 3 / 1
15/21
claude-opus-5
Claude Code 2.1.278
21
18
13… See the full description on the dataset page: https://huggingface.co/datasets/yangjunxiao2021/mle-trajectories.machine-failure-mlops-demo-logsml-intern-sessions
ML Intern session traces
This dataset contains ML Intern coding agent session traces uploaded from local
ML Intern runs. The traces are stored as JSON Lines files under sessions/,
with one file per session.
Links
ML Intern demo: https://smolagents-ml-intern.hf.space
ML Intern CLI: https://github.com/huggingface/ml-intern
Data description
Each *.jsonl file contains a single ML Intern session converted to a
Claude-Code-style event stream for the… See the full description on the dataset page: https://huggingface.co/datasets/merve/ml-intern-sessions.MLR_sft_data
MLR SFT Data
MLR SFT Data is a teacher-generated supervised fine-tuning dataset for training Multi-Level Reasoning (MLR) models in the paper Enhancing Language Model Reasoning with Structured Multi-Level Modeling (ICLR 2026). It decomposes complete reasoning trajectories into two types of step-level examples:
Planner: plans the next reasoning goal and task based on the problem and reasoning history.
Executor: executes the Planner's instruction and updates the reasoning state.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_sft_data.whestbench-relu-mlp-moments-10k
WhestBench Random ReLU MLPs — Monte-Carlo Activation Cumulants (10k)
A dataset of 10,500 random ReLU MLPs (10,000 train + 500 held-out test) together with
Monte-Carlo–estimated per-layer activation cumulants (mean, variance, skewness, kurtosis) for
every layer, for both the pre-activation and post-ReLU signals. Built for the WhestBench
Estimation Challenge 2026 and for research on analytic moment / uncertainty propagation
through deep networks.
The generative process… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-moments-10k.traces.claude-code.mlx-lm-granitemoehybridswe_v0.1_jsonl_wo_mlang_large100_wo_v0.0_deltagMLR_structured_trajectory
Reasoning Trajectories with Step-Level Annotations
This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling.
Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.MLR_full_trajectory
DeepSeek-R1 Reasoning Trajectories
This dataset contains raw reasoning trajectories generated by DeepSeek-R1, used in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling.
The trajectories capture the full reasoning process produced by the model, including hidden chain-of-thought reasoning and final responses.
They are provided to support research on reasoning analysis, trajectory supervision, and multi-step reasoning training.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_full_trajectory.ovos-stt-bench-mls-es-ES
OVOS stt bench — mls-es-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-es-ES.ML_MTCONAN_KN
1st Workshop on Multilingual Counterspeech generation: Shared Task
Dataset description
The datasets consist of 596 Hate Speech-Counter Narrative pairs. In this dataset, the hate speech is taken from MTCONAN, while the counter narratives are newly generated. Together with each pair, we also provide 5 background knowledge sentences, some of which are relevant for obtaining the counter narratives. The dataset is available in 4 different languages (Basque, English… See the full description on the dataset page: https://huggingface.co/datasets/LanD-FBK/ML_MTCONAN_KN.MedHorizon
MedHorizon
MedHorizon is a long-context medical video benchmark for evaluating multimodal models on full-procedure clinical videos. The benchmark emphasizes two properties that are not captured by short-clip medical video datasets: extremely sparse evidence retrieval and multi-hop reasoning over observations distributed across a full procedure.
Dataset Contents
Videos: 340 full-procedure videos.
Questions: 1,253 multiple-choice QA pairs.
Evaluation split: test.
Video… See the full description on the dataset page: https://huggingface.co/datasets/mlvbench-review/MedHorizon.ml-intern-sessions
ML Intern session traces
This dataset contains ML Intern coding agent session traces uploaded from local
ML Intern runs. The traces are stored as JSON Lines files under sessions/,
with one file per session.
Links
ML Intern demo: https://smolagents-ml-intern.hf.space
ML Intern CLI: https://github.com/huggingface/ml-intern
Data description
Each *.jsonl file contains a single ML Intern session converted to a
Claude-Code-style event stream for the… See the full description on the dataset page: https://huggingface.co/datasets/xucenying/ml-intern-sessions.lm-eval-results-mlabonne-Zebrafish-7B-private
Dataset Card for Evaluation run of mlabonne/Zebrafish-7B
Dataset automatically created during the evaluation run of model mlabonne/Zebrafish-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-mlabonne-Zebrafish-7B-private.ovos-stt-bench-mls-pt-PT
OVOS stt bench — mls-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.ovos-stt-bench-mls-fr-FR
OVOS stt bench — mls-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-fr-FR.
