datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.VisionThink-General-Train
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao/VisionThink-General-Train
This is the training dataset used for our Reasoning VLM on general VQA tasks.
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper]
Senqiao Yang,
Junyi Li,
Xin Lai,
Bei Yu,
Hengshuang Zhao,
Jiaya Jia
Highlights
Our VisionThink leverages reinforcement learning to autonomously learn whether to… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-General-Train.Valen-Training-General-100k
Valen-Training-General-100k
GitHub · Preview model · General evaluation · Technical notes
✨ Introduction
Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 100,000 image-based decision records for supervised training and decision-head learning in Valen, spanning visual question answering, interfaces, games and documents.
Each record contains one decision question… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Training-General-100k.piperx-general-pickup-clutter-preview
Piper X General Pickup — Same-cluster preview
▶ Play episode 0
4 successful episodes, 694 frames, 25 FPS; three synchronized 640×480 RGB cameras.
这一版将目标和干扰物放在同一簇内,取消干扰物的前后分区以及目标周围 7.5 cm 的统一禁入区。
至少三个干扰物初始摆在目标 XY 包围盒附近,初始间距为 1.5–3 cm;其余物体沿同一簇扩展。
物理稳定后与录制开始时,均要求至少三个刚体干扰物到目标的 XY 包围盒间距不超过 6 cm。
这些距离为保守包围盒足迹代理,不能视为精确网格表面距离或遮挡指标。
本预览为近邻、不重叠的桌面布局;不包含有意堆叠。
The target is physically grasped near its live mesh centre, approached and descended to with open fingers,
then closed upon only… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-general-pickup-clutter-preview.Valen-Eval-General-5k
Valen-Eval-General-5k
GitHub · 中文 README · Preview model · Technical notes
✨ Introduction
Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 5,000 image-based decision records for held-out evaluation, spanning visual question answering, interfaces, games and documents.
Each record contains one decision question, a target probability distribution, local image… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Eval-General-5k.Project-Imaging-X
Project Imaging-X is a strategic initiative to consolidate 1000+ open medical imaging datasets worldwide, breaking down data silos through systematic integration to build the foundational infrastructure for next-generation medical AI models.
Challenge: Medical imaging lacks large-scale unified datasets due to clinical expertise requirements and privacy constraints, limiting the development of powerful medical foundation models.
Solution: We surveyed 1000+… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/Project-Imaging-X.General_pic_20230326general_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetmarqo-general-ecommerce-evalshowdown-clicks
showdown-clicks
General Agents
🤗 Dataset | GitHub
showdown is a suite of offline and online benchmarks for computer-use agents.
showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents.
As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.google-shopping-general-eval
Marqo Ecommerce Embedding Models
In this work, we introduce the GoogleShopping-1m dataset for evaluation. This dataset comes with the release of our state-of-the-art embedding models for ecommerce products: Marqo-Ecommerce-B and Marqo-Ecommerce-L.
Released Content:
Marqo-Ecommerce-B and Marqo-Ecommerce-L embedding models
GoogleShopping-1m and AmazonProducts-3m for evaluation
Evaluation Code
The benchmarking results show that the… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/google-shopping-general-eval.General-Evol-VQA
Dataset Card for General-Evol-VQA-1.2M
This dataset has been carefully curated to enhance the general instruction capabilities of Vision-Language Models (VLMs). It comprises two subsets:
600k English samples
600k Korean samples
We recommend using this dataset alongside other task-specific datasets (e.g., OCR, Language, code, math, ...) to improve performance and achieve more robust model capabilities.
Made by: maum.ai Brain NLP. Jaeyoon Jung, Yoonshik Kim
Dataset Target… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/General-Evol-VQA.GeneralScience-MLLM-22K
GeneralScience-MLLM-22K
Dataset Summary
GeneralScience-MLLM-22K is a unified general-science multiple-choice QA collection built from local snapshots of SciQ, AI2 ARC, and ScienceQA. It follows the same release style as a subject-specific MLLM dataset: every sample is stored as one JSONL record, text-only and image-text examples share one schema, and ScienceQA images are exported as standalone files referenced by relative paths.
The release contains 22,661… See the full description on the dataset page: https://huggingface.co/datasets/gineven/GeneralScience-MLLM-22K.IMed-361M
IMed-361M
🤝 免责声明
IMed-361M是由多个公开的数据集组成,旨在取之于社区,回馈于社区,为研究人员和开发者提供一个用于学术和技术研究的资源。使用本数据集的任何个人或组织(以下统称为“使用者”)需遵守以下免责声明:
数据集来源:本数据集由多个公开的数据集组成,这些数据集的来源已在预印版论文中明确标明。使用者应当遵守原始数据集的相关许可和使用条款。
数据准确性:尽管我们已经努力确保数据集的准确性和完整性,但无法对数据集的准确性作出保证。使用者应自行承担使用数据集可能带来的风险和责任。
责任限制:在任何情况下,数据集的提供者及相关贡献者均不对使用者的任何行为或结果承担责任。
使用约束:使用者在使用本数据集时,应遵守适用的法律法规和伦理规范。使用者不得将本数据集用于非法、侵犯隐私、诽谤、歧视或其他违法或不道德的目的。
知识产权:本数据集的知识产权归原始数据集的相关权利人所有,使用者不得以任何方式侵犯数据集的知识产权。… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/IMed-361M.general_20_data_remove_v0SEC_Tables_Lite
Dataset Card for "SEC_Tables_Lite"
More Information needed
General_DatasetMirror or HerdNet General Dataset to make the HerdNet Code implementation easier to reproduce
Delplanque, Alexandre; Foucher, Samuel; Lejeune, Philippe; Linchant, Julie; Théau, Jérôme, 2023, "Dataset & Code for paper: "Multispecies detection and identification of African mammals in aerial imagery using convolutional neural networks"", https://doi.org/10.58119/ULG/MIRUU5, ULiège Open Data Repository, V1, UNF:6:cV/uw5pGcUHzctn8ROmotg== [fileUNF]… See the full description on the dataset page: https://huggingface.co/datasets/karisu/General_Dataset.google-shopping-general-eval-100k
Marqo Ecommerce Embedding Models
In this work, we introduce the GoogleShopping-1m dataset for evaluation. This dataset comes with the release of our state-of-the-art embedding models for ecommerce products: Marqo-Ecommerce-B and Marqo-Ecommerce-L.
Released Content:
Marqo-Ecommerce-B and Marqo-Ecommerce-L embedding models
GoogleShopping-1m and AmazonProducts-3m for evaluation
Evaluation Code
The benchmarking results show that the… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/google-shopping-general-eval-100k.general-pmd-synthetic-testing-with-embeddingsThis dataset is designed to be used in testing. It's derived from general-pmd-10k datasetThumbnailTrendingDatasetsaraiki-ocr-general-dataset
Saraiki OCR General Dataset ( Jataki dialect)
generalized-ablations-resultslocalize-general-objects
Elliot Localize General Objects
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. Coordinates are normalized 0–1000. Within each query targets are sorted left-to-right then top-to-bottom.
HF preview configs contain 10 examples per view, not the complete training split. Full training… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-general-objects.general_cv_dataset_6general_cv_dataset_8general-off2on-filteredbcgeneral_cv_dataset_7factures_generales
Dataset Card for "factures_generales"
More Information needed
ThumbnailDataset
