datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.RADAR-auxiliary-data
RADAR: Preprocessed Anatomical Masks for Merlin CT Data
This dataset provides preprocessed anatomical segmentation masks for the Merlin abdominal CT training set, generated by TotalSegmentator and post-processed for use with the RADAR framework. These masks enable anatomy-aware vision–language pretraining without any additional manual annotation.
Overview
RADAR is a generalist vision–language model trained on over 400,000 contrast-enhanced abdominal CT… See the full description on the dataset page: https://huggingface.co/datasets/radar-generalist/RADAR-auxiliary-data.SlideChat
Introduction
This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding.
The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks.
Contents
Training Instruction Data
SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.General-Bench-Closeset
On Path to Multimodal Generalist: General-Level and General-Bench
[📖 Project]
[🏆 Leaderboard]
[📄 Paper]
[🤗 Paper-HF]
[🤗 Dataset-HF]
[📝 Dataset-Github]
Close Set of General-Bench
We divide our General-Bench into two settings: open and close.
This is the Close Set, where we release only the sample inputs—without ground-truth answers—for 🏆 Leaderboard purpose.
To participate the leaderboard, please follow the detailed instructions to submit the evaluation results (submission).… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Closeset.generalization-science-datashofo-tiktok-general-small
Shofo TikTok General (Small)
Overview
Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos.
Size: ~50K videos (~500GB)
Modality: Video + Audio + Text (transcripts, comments, captions)
Source: TikTok
Schema
Column
Type
Description
file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.MiMo-V2.6-RL-harbor-general
MiMo-V2.6-RL General (Harbor)
Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's
MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained
on, converted so every one runs as a standard Harbor task.
Each task is a workplace with 5 to 15 business systems (finance, legal, HR, operations, ...) served over MCP, a workspace of documents, and a brief. The agent works through the systems as an unprivileged user; the task's own… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-general.Ophora-160K
Introduction
Ophora-160K contains 162,185 video clip-instruction pairs extracted from 9,819 narrative videos of ophthalmic surgery. The average duration of all clips is 5.54 seconds.
All ophthalmic narrative videos were collected from the YouTube platform. The file ophora161k.csv contains all video clip IDs and their generation instructions. The file ophora28k.csv contains content that has been filtered to remove sensitive information such as subtitles and watermarks.… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/Ophora-160K.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.steady-rans-generalization
Steady-RANS cross-family generalization dataset
Data for the paper "Towards generalized flow field prediction: one model across unseen
object families" (under double blind review; this account is anonymous for that reason).
Trained checkpoints and evaluation code are in the companion model repo:
steady-rans-surrogates.
Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct
shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data
General-Bench-Openset
On Path to Multimodal Generalist: General-Level and General-Bench
[📖 Project]
[🏆 Leaderboard]
[📄 Paper]
[🤗 Paper-HF]
[🤗 Dataset-HF (Close-Set)]
[🤗 Dataset-HF (Open-Set)]
[📝 Github]
Open Set of General-Bench
We divide our General-Bench into two settings: Open and Close.
This is the Open Set repo, where we release the full ground-truth annotations for all datasets, allowing to train and evaluate models for open research purpose.
If you wish to rank on our 🏆 leaderboard, please… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Openset.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.math-code-general-prunedOpenOneRec-General-Pretrain
通用文本数据集
本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。
数据格式说明
所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持:
Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表
Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表
每个 Parquet 文件包含以下核心字段:
uuid: 唯一标识符
source: 数据来源标识
metadata: JSON 格式的元数据字典
segments 或 messages: 文本内容(根据数据类型选择)
详细的数据格式规范请参考 ../README.md。
数据集列表
数据集名称
样本数量
HuggingFace 仓库
reasoning_v1_20m
1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.HyperWorld-Bench-GeneralHelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-chat-formatted-generationsVisionThink-General-Train
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao/VisionThink-General-Train
This is the training dataset used for our Reasoning VLM on general VQA tasks.
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper]
Senqiao Yang,
Junyi Li,
Xin Lai,
Bei Yu,
Hengshuang Zhao,
Jiaya Jia
Highlights
Our VisionThink leverages reinforcement learning to autonomously learn whether to… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-General-Train.GeneralThoughtArchive
GeneralThought-430K
Thought wants to be free
Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this
dataset but are archiving it here.
The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.AITW_GeneralValen-Training-General-100k
Valen-Training-General-100k
GitHub · Preview model · General evaluation · Technical notes
✨ Introduction
Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 100,000 image-based decision records for supervised training and decision-head learning in Valen, spanning visual question answering, interfaces, games and documents.
Each record contains one decision question… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Training-General-100k.general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.IMG_GENERAL
Usage
This dataset is provided for research purposes.
Redistribution, mirroring, or republishing of the dataset files
is not permitted.
The uploader does not claim ownership of the underlying content.
Users are responsible for complying with applicable copyright
and other rights.
piperx-general-pickup
Piper X General Pickup
已发布 79 个资产、3753/4420 条成功 episode。
每个资产目标 52 条成功轨迹;成功率过低、尝试预算内未能采满的资产按实际成功条数发布。
采集状态(截至 2026-10-08)
采集已结束。计划 85 个资产、每个 52 条,共 4,420 条;实际情况:
已采集 79 个资产、3,753 条(约 85%):69 个采满 52 条,10 个不满额,按实际成功条数发布。
未采集 6 个资产:在尝试预算内 0 条成功,数据集中没有它们的 episode。
不满额资产:speaker/00000 42、trophy/00000 10、garage/00017 7、bell/00000 5、corkscrew/00000 9、bottle_opener/00000 2、cactus/00000 4、chocolate_bag/00000 36、pot/00001 37、wrench/00002 13。
未采集资产:
资产
尝试次数
情况… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-general-pickup.piperx-general-pickup-clutter-preview
Piper X General Pickup — Same-cluster preview
▶ Play episode 0
4 successful episodes, 694 frames, 25 FPS; three synchronized 640×480 RGB cameras.
这一版将目标和干扰物放在同一簇内,取消干扰物的前后分区以及目标周围 7.5 cm 的统一禁入区。
至少三个干扰物初始摆在目标 XY 包围盒附近,初始间距为 1.5–3 cm;其余物体沿同一簇扩展。
物理稳定后与录制开始时,均要求至少三个刚体干扰物到目标的 XY 包围盒间距不超过 6 cm。
这些距离为保守包围盒足迹代理,不能视为精确网格表面距离或遮挡指标。
本预览为近邻、不重叠的桌面布局;不包含有意堆叠。
The target is physically grasped near its live mesh centre, approached and descended to with open fingers,
then closed upon only… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-general-pickup-clutter-preview.Valen-Eval-General-5k
Valen-Eval-General-5k
GitHub · 中文 README · Preview model · Technical notes
✨ Introduction
Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 5,000 image-based decision records for held-out evaluation, spanning visual question answering, interfaces, games and documents.
Each record contains one decision question, a target probability distribution, local image… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Eval-General-5k.theo_impulsive-generality-bank-lineage
Status: NOT the paper's results. Build lineage (archive, generated, generated_v2) of the impulsive generality eval banks: eval inputs, not results. The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
theo_impulsive-generality-bank-lineage
Objaverse-General-Find3DThis dataset contains benchmarks for open-world object part segmentation proposed by Find Any Part in 3D (ICCV 2025).
This dataset includes two human-annotated benchmarks: Objaverse-General (of 100 object categories) and ShapeNetPart-Objaverse (of the same categories of ShapeNetPart, but with objects source from Objaverse to study distribution shift).
Usage
Inside both objaverse-general and objaverse-shapanetepart directories, the benchmark has the following directory structure:… See the full description on the dataset page: https://huggingface.co/datasets/ziqima/Objaverse-General-Find3D.standard_chat_tool_calling_general
