datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.VQAonline
VQAonline
🌐 Homepage | 🤗 Dataset | 📖 arXiv
Dataset Description
We introduce VQAonline, the first VQA dataset in which all contents originate from an authentic use case.
VQAonline includes 64K visual questions sourced from an online question answering community (i.e., StackExchange).
It differs from prior datasets; examples include that it contains:
(1) authentic context that clarifies the question
(2) an answer the individual asking the question validated as… See the full description on the dataset page: https://huggingface.co/datasets/ChongyanChen/VQAonline.ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.MVBench
MVBench
Important Update
[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference.
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.M3_VOS
[CVPR 2025] M3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Venue: CVPR2025
Repository: 🛠️Tool, 🏠Page
Paper: arxiv.org/html/2412.13803v2
Point of Contact: Jiaxin Li , Zixuan Chen
📁 Structure
This dataset contains annotated videos and images for object segmentation tasks with phase transition information. The directory… See the full description on the dataset page: https://huggingface.co/datasets/Lijiaxin0111/M3_VOS.ExtractBench
ExtractBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.LongDS
LongDS
Overview •
Versions •
Download •
Usage •
Task Index •
Citation
Paper •
Code •
Web •
📊 Leaderboard
LongDS-Bench is a benchmark for evaluating long-horizon, multi-turn agentic data analysis. Real-world analysis is rarely a sequence of independent questions: filters, metric definitions, assumptions, intermediate tables, and branch-specific results evolve over many turns. LongDS tests whether agents can maintain and apply these… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/LongDS.FormStruct-Bench
FormStruct-Bench
Dataset Description
FormStruct-Bench is a multilingual benchmark for extracting the semantic and
spatial structure of forms from document images. The repository combines a
7,000-page main benchmark, a controlled visual-degradation set, and
template-level layout annotations. It supports evaluation of vision-language
models and document AI systems on hierarchical key-value extraction, document
structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.GUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.hateful_memes_expandedFinch
Finch (FinWorkBench): Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
This repository contains the dataset for Finch, an enterprise-grade benchmark for evaluating an agent’s ability to work like a skilled finance & accounting expert (work IQ) on real-world professional workflows.
Paper: https://arxiv.org/abs/2512.13168
Evaluation Code: https://github.com/FinWorkBench/Finch
🍻 Updates
2026-4-6: FinWorkBench is accepted to ACL… See the full description on the dataset page: https://huggingface.co/datasets/FinWorkBench/Finch.PKLotPKLot: Parking Space Occupancy Detection Dataset
Unofficial redistribution of the PKLot parking-lot database, reformatted into a standardized YOLO-compatible detection layout, under the original CC BY 4.0 license.
Disclaimer
This repository is not an official release of the PKLot dataset.
PKLot was created by Paulo R. L. de Almeida, Luiz S. Oliveira, Alceu S. Britto Jr, Eunelson J. Silva Jr, and Alessandro L. Koerich, and is distributed by the Vision, Robotics and… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/PKLot.ChartDiff
ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts
Overview
ChartDiff is a large-scale benchmark for cross-chart comparative summarization, designed to evaluate whether vision-language models can identify differences and generate coherent comparative descriptions across pairs of charts.
Unlike existing chart understanding datasets that emphasize single-chart interpretation, ChartDiff requires models to compare two charts jointly and generate a concise… See the full description on the dataset page: https://huggingface.co/datasets/ckchaos/ChartDiff.SeaShipsSeaShips (7000): Maritime Ship Detection Dataset
Unofficial redistribution of the SeaShips(7000) maritime ship-detection dataset, reformatted into a standardized YOLO-compatible directory layout. License status is unclear -- see License before using this beyond research.
Disclaimer
This repository is not an official release of SeaShips.
SeaShips was created by Zhenfeng Shao, Wenjing Wu, Zhongyuan Wang, Wan Du, and Chengyuan Li at the State Key Laboratory of… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/SeaShips.ShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.jev-bench
jev-bench
Real human-labeled data, reformatted into System One questions — with human label distributions wherever they exist.
22 configs · 166,054 rows · 22,773 test records · 4 calibration-gold configs · 46 models scored · v0.1.1
Code & engine · Findings · Leaderboard · Published models · Source rationale · Jev's API, verified · Other Jev evaluations
Every model on the same 22,773 test records. Down and to the right is better; the stars are open models fine-tuned on their… See the full description on the dataset page: https://huggingface.co/datasets/Praveenrajus/jev-bench.MedXpertQA
Dataset Card for MedXpertQA
MedXpertQA is a highly challenging and comprehensive benchmark designed to evaluate expert-level medical knowledge and advanced reasoning capabilities. It features both text-based and multimodal question-answering tasks, with the multimodal subset leveraging structured clinical information alongside images.
Dataset Description
MedXpertQA comprises 4,460 questions spanning diverse medical specialties, tasks, body systems, and image types. It… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/MedXpertQA.VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.flat-pack-bench
Flat-Pack Bench 🧩
Furniture assembly as a spatio-temporal stress test for large vision-language models.
Flat-Pack Bench is a multiple-choice benchmark for evaluating fine-grained
spatio-temporal understanding in real furniture assembly videos. Each question
asks a model to reason about object parts, contact events, assembly order, final
connectivity, or part identity across time.
Project page: https://flat-pack-bench.github.io
🎯 Benchmark Tasks
The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/justachetan/flat-pack-bench.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.DBIndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.svgrepo
Dataset Card for SVGRepo Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGRepo.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under a specific open-source or permissive license, clearly indicated in its metadata. The SVG… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgrepo.turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.vstar_benchSLAKE
Dataset Info:
SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering [ISBI 2021 oral]
Project Page: click
Corresponding Authors: Bo Liu, Xiao-Ming Wu
Any questions, please contact us. Thank you!
Modification:
In the Huggingface Repo, we have changed the name of validate.json to validation.json to better display in the Dataset Card.
