Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clip-benchmark /wds_objectnetimage1K<n<10K4 likes81k downloads4y agoHugging Face02bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes66k downloads1y agoHugging Face03benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes24k downloads5mo agoHugging Face04clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes20k downloads4y agoHugging Face05BLINK-Benchmark /BLINK BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.image1K<n<10K49 likes13k downloads1y agoHugging Face06kohsei /MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌 CVPR 2026 (Main) This repository provides the datasets for “MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta Paper Link https://arxiv.org/abs/2511.22989 Github Repository For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.imagetext-to-image1K<n<10K5 likes12k downloads3mo agoHugging Face07clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes12k downloads4y agoHugging Face08gaia-benchmark /GAIAgated GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.audion<1K880 likes11k downloads1y agoHugging Face09clip-benchmark /wds_imagenet-aimage1K<n<10K0 likes11k downloads4y agoHugging Face10SLM-Lab /benchmark SLM Lab Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book Foundations of Deep Reinforcement Learning. Documentation · Benchmark Results NOTE: v5.0 updates to Gymnasium, uv tooling, and modern dependencies with ARM support - see CHANGELOG.md. Book readers: git checkout v4.1.1 for Foundations of Deep Reinforcement Learning code. BeamRider Breakout KungFuMaster MsPacman Pong Qbert Seaquest Sp.Invaders… See the full description on the dataset page: https://huggingface.co/datasets/SLM-Lab/benchmark.image1K<n<10K0 likes10k downloads7mo agoHugging Face11clip-benchmark /wds_imagenet1kimage10K<n<100K1 likes10k downloads4y agoHugging Face12Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M37 likes9.9k downloads1mo agoHugging Face13DabbyOWL /PDE_Inverse_Problem_Benchmarking PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems. Code: GitHub - ASK-Berkeley/PDEInvBench Sample Usage You can use the provided script from the codebase to batch download the data: pip install huggingface_hub python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.imageother100M<n<1B3 likes8.6k downloads5mo agoHugging Face14clip-benchmark /wds_imagenetv2image10K<n<100K0 likes8.1k downloads4y agoHugging Face15dpdl-benchmark /oxford_flowers102image1K<n<10K12 likes6.5k downloads2y agoHugging Face16RoboDojo-Benchmark /GOAI-2026imagen<1K0 likes4.5k downloads1mo agoHugging Face17clip-benchmark /wds_fer2013image10K<n<100K0 likes4k downloads4y agoHugging Face18zhiyuzhang-0212 /MOVA_benchmark_for_arena MOVA Benchmark for Arena This is the benchmark used for the subjective arena experiments of MOVA (MOVA: Towards Scalable and Synchronized Video–Audio Generation). All prompts are rewritten by the workflow introduced in the paper. Paper: MOVA: Towards Scalable and Synchronized Video–Audio Generation Code: https://github.com/OpenMOVA/MOVA Overview The benchmark contains 732 samples in total, organized into two subsets: Subset Samples MOVA-Bench 132… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuzhang-0212/MOVA_benchmark_for_arena.image1K<n<10K3 likes3.7k downloads6mo agoHugging Face19issai /Food_Portion_Benchmark Food Portion Benchmark (FPB) Dataset The Food Portion Benchmark (FPB) is a comprehensive dataset and benchmark suite for multi-task food scene understanding, combining food detection and portion size (weight) estimation. It was introduced to support research in dietary analysis, nutrition tracking, and food computing. The dataset is built with high-quality annotations and evaluated using an extended YOLOv12-based multi-task model . 📦 Dataset Overview Total images:… See the full description on the dataset page: https://huggingface.co/datasets/issai/Food_Portion_Benchmark.image8 likes3.5k downloads11mo agoHugging Face20tccoin /navverse-benchmark NavVerse Benchmark This repository hosts the NavVerse dataset. Runnable scene release navverse_v1_july23.tar.gz is the current runnable scene package. It contains: 52 VC+ outdoor scenes and 52 connected indoor/outdoor scenes under vc_plus/; 30 GRScenes commercial navigation scenes under grscenes_commercial/; the corresponding prebuilt navmesh/ assets; a relative nvidia -> vc_plus/nvidia link for the bundled CloudySky runtime lighting. From a NavVerse-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/tccoin/navverse-benchmark.imagen<1K1 likes3.3k downloads2mo agoHugging Face21LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K2 likes3.2k downloads1y agoHugging Face22IDEAL-Benchmark /IDEAL-Scenes IDEAL-Bench: Indoor Dataset for Evaluating Analysis by 3D Layout Reasoning IDEAL-Bench is an evaluation suite that requires VLMs to predict structured 3D layouts on photorealistic indoor scenes across 10 room types, scored along five numerical dimensions (scene validity, physical plausibility, geometric accuracy, object recognition, and grid layout) and a perceptual render-and-compare protocol. Built on IDEAL-Scenes - 1,000 procedurally generated, re-renderable Blender scenes… See the full description on the dataset page: https://huggingface.co/datasets/IDEAL-Benchmark/IDEAL-Scenes.imageimage-to-text1K<n<10K0 likes2.8k downloads3mo agoHugging Face23FanqingM /MMIU-Benchmark Dataset Card for MMIU Repository: https://github.com/OpenGVLab/MMIU Paper: https://arxiv.org/abs/2408.02718 Project Page: https://mmiu-bench.github.io/ Point of Contact: Fanqing Meng Introduction MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular MLLMs, including both open-source and proprietary models… See the full description on the dataset page: https://huggingface.co/datasets/FanqingM/MMIU-Benchmark.image10K<n<100K11 likes2.6k downloads2y agoHugging Face24billhdzhao /FedRemoteSensing_Benchmark Dataset README 1. General Information Number of Labels: There are a total of 5 labels, namely: Agriculture, Bareland, Forest, Residential, and River. Number of Clients: The dataset consists of 100 clients. Data Volume per Client: Each client contains approximately 350 tif format images. 2. Data Sources All the images are collected from 6 different datasets, which are as follows: Eurosat UC Merced Land Use Dataset AID NWPU - RESISC45 WHU-RS19 NaSC-tg2 The data… See the full description on the dataset page: https://huggingface.co/datasets/billhdzhao/FedRemoteSensing_Benchmark.image10K<n<100K2 likes2.4k downloads2y agoHugging Face25BloomBerry /figma-slide-benchmark Figma Slide Editing Benchmark Benchmark accompanying our EMNLP 2026 Industry Track (Main) accepted paper "ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation". 📄 Paper: https://arxiv.org/pdf/2608.24103 💻 Code: https://github.com/BloomBerry/agentic-canvas-editor Overview Each benchmark item is a slide-editing task defined as a pair of Figma Slides documents: *_TestA — the input deck the agent starts from. *_GroundTruthA — the… See the full description on the dataset page: https://huggingface.co/datasets/BloomBerry/figma-slide-benchmark.image1K<n<10K1 likes2.4k downloads11d agoHugging Face26NL3D /NL3D-Synth-Benchmark Dataset Card for NL3D-Synth-Benchmark Dataset Summary NL3D-Synth-Benchmark is the synthetic evaluation split of NL3D, a synthetic-real dataset and benchmark for non-Lambertian 3D reconstruction. This split is designed to provide a compact, controlled, and reproducible benchmark for evaluating 3D perception methods on challenging non-Lambertian materials. The benchmark contains 16 synthetic scenes rendered with physically based material models and full 3D background… See the full description on the dataset page: https://huggingface.co/datasets/NL3D/NL3D-Synth-Benchmark.image10K<n<100K0 likes2.3k downloads5mo agoHugging Face27vanthanh /UAVDT-Benchmark-Mimage10K<n<100K0 likes2.3k downloads3mo agoHugging Face28drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face29morzel85 /synthetic-medical-document-recognition-benchmark Synthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the dataset suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.documentimage-to-text10K<n<100K1 likes2.1k downloads6d agoHugging Face30brennercruvinel /mtg-urna-benchmark 38,627 Magic: The Gathering cards, one per oracle id, the scan and the rules text of each, packed into single .urna files that answer text and image queries from memory-mapped bytes: no server, no Python at read time. Ten such files live here. They carry the same cards, the same text and the same content hash; what differs is how the 4 GB of JPEG was encoded inside, and which image models embedded it. The file to start with is release/v0.3/stills-5models. It is the only one with the models… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/mtg-urna-benchmark.imageimage-to-text100K<n<1M2 likes2k downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.