Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Nithish2410 /benchmark-bcplustextn<1K0 likes115k downloads7mo agoHugging Face02clip-benchmark /wds_objectnetimage1K<n<10K4 likes72k downloads4y agoHugging Face03bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes63k downloads1y agoHugging Face04benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes22k downloads5mo agoHugging Face05clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes20k downloads4y agoHugging Face06alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes19k downloads18h agoHugging Face07kohsei /MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌 CVPR 2026 (Main) This repository provides the datasets for “MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta Paper Link https://arxiv.org/abs/2511.22989 Github Repository For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.imagetext-to-image1K<n<10K5 likes17k downloads3mo agoHugging Face08AlphaDojo /dojo_benchmark_kline Languages: 简体中文 · English dojo_benchmark_kline — Benchmark Index Bars Overview Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date. Files File Description data.parquet Index daily bars Key Fields Field Description symbol Index code (e.g. ^SPX, 000300.SS) kline_t Bar interval; "1D" for daily bars bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.text1K<n<10K1 likes16k downloads6h agoHugging Face09google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K267 likes15k downloads2y agoHugging Face10BLINK-Benchmark /BLINK BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.image1K<n<10K49 likes14k downloads1y agoHugging Face11MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K49 likes13k downloads2mo agoHugging Face12leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes12k downloads6mo agoHugging Face13simular-ai /sai-osworld-v2-benchmark-runs Sai on OSWorld-V2 — benchmark runs of record Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API protocol logs, and run manifests. Run Date Tasks scored Mean score Perfect (1.0) Zeros run1/ 2026-08-12 108/108 0.7276 28 7 run2/ 2026-08-20 108/108 0.7329 33 5 Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.textother2 likes11k downloads1mo agoHugging Face14gaia-benchmark /GAIAgated GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.audion<1K865 likes11k downloads11mo agoHugging Face15clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes11k downloads4y agoHugging Face16UMIbenchmark /UMI-Benchmark-v1 UMI-Benchmark-v1 UMI-Benchmark-v1 contains 20,000 real-world robot manipulation sessions across 10 tasks: 4 single-arm tasks and 6 dual-arm tasks. Dataset Summary Category Tasks Sessions Single-arm 4 9,988 Dual-arm 6 10,012 Total 10 20,000 Tasks ID Folder Type Task Sessions T1 single_arm_task1 Single-arm Stack Baskets 3,000 T2 single_arm_task2 Single-arm Trash Bag 1,991 T3 single_arm_task3 Single-arm Stamp Ink 2… See the full description on the dataset page: https://huggingface.co/datasets/UMIbenchmark/UMI-Benchmark-v1.text10K<n<100K3 likes10k downloads8d agoHugging Face17Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M36 likes9.8k downloads1mo agoHugging Face18clip-benchmark /wds_imagenet-aimage1K<n<10K0 likes9.4k downloads4y agoHugging Face19nlile /hendrycks-MATH-benchmark Hendrycks MATH Dataset Dataset Description The MATH dataset is a collection of mathematics competition problems designed to evaluate mathematical reasoning and problem-solving capabilities in computational systems. Containing 12,500 high school competition-level mathematics problems, this dataset is notable for including detailed step-by-step solutions alongside each problem. Dataset Summary The dataset consists of mathematics problems spanning multiple… See the full description on the dataset page: https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark.text10K<n<100K33 likes9.3k downloads2y agoHugging Face20madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M0 likes8.8k downloads26d agoHugging Face21clip-benchmark /wds_imagenetv2image10K<n<100K0 likes7.6k downloads4y agoHugging Face22clip-benchmark /wds_imagenet1kimage10K<n<100K1 likes7.5k downloads4y agoHugging Face23YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes7.2k downloads1mo agoHugging Face24pdfqa /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.documentquestion-answering5 likes7k downloads7mo agoHugging Face25bop-benchmark /megaposetext1M<n<10M1 likes6.6k downloads2y agoHugging Face26PRHW /loom-benchmark-mmlu-protext10K<n<100K0 likes6.3k downloads3mo agoHugging Face27md-nishat-008 /mHumanEval-Benchmark 🔷 Accepted in NAACL Proceedings (2025) 🔷 mHumanEval The mHumanEval benchmark is curated based on prompts from the original HumanEval 📚 [Chen et al., 2021]. It includes a total of 33,456 prompts for Python, and 836,400 in total - significantly expanding from the original 164. Quick Start Detailed… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/mHumanEval-Benchmark.text100K<n<1M4 likes6.1k downloads1y agoHugging Face28BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes5.8k downloads1y agoHugging Face29inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes5.6k downloads4mo agoHugging Face30nvidia /cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset. Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes. This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub. textn<1K40 likes4.5k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.