Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RoboDojo-Benchmark /RoboDojo10 likes187k downloads7d agoHugging Face02Nithish2410 /benchmark-bcplustextn<1K0 likes115k downloads7mo agoHugging Face03DL3DV /DL3DV-Benchmarkgated DL3DV Benchmark Download Instructions This repo contains 140 scenes in the DL3DV-benchmark, which are sampled from DL3DV-10K. The repo includes a README, License, colmaps/images (compatible to nerfstudio and 3D gaussian splatting), scene labels and the performances of methods reported in the paper (ZipNeRF, 3DGS, MipNeRF-360, nerfacto, Instant-NGP). The benchmark preview page can be found here https://dl3dv-10k.github.io/DL3DV-Benchmark-Preview/. Download As the whole… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Benchmark.n>1T47 likes73k downloads1y agoHugging Face04clip-benchmark /wds_objectnetimage1K<n<10K4 likes72k downloads4y agoHugging Face05bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes63k downloads1y agoHugging Face06hf-benchmarks /transformers2 likes27k downloads2h agoHugging Face07benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes22k downloads5mo agoHugging Face08clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes20k downloads4y agoHugging Face09alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K5 likes19k downloads14h agoHugging Face10vedangfake /chess-slm-benchmark0 likes18k downloads28m agoHugging Face11kohsei /MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌 CVPR 2026 (Main) This repository provides the datasets for “MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta Paper Link https://arxiv.org/abs/2511.22989 Github Repository For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.imagetext-to-image1K<n<10K5 likes17k downloads3mo agoHugging Face12AlphaDojo /dojo_benchmark_kline Languages: 简体中文 · English dojo_benchmark_kline — Benchmark Index Bars Overview Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date. Files File Description data.parquet Index daily bars Key Fields Field Description symbol Index code (e.g. ^SPX, 000300.SS) kline_t Bar interval; "1D" for daily bars bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.text1K<n<10K1 likes16k downloads2h agoHugging Face13google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K267 likes15k downloads2y agoHugging Face14mackelab /benchmarking_sbi_runs Benchmarking SBI Runs This dataset contains the raw, per-run results underlying the manuscript "Benchmarking Simulation-Based Inference" (Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021). It is a direct migration of the Git LFS data from mackelab/benchmarking_sbi_runs on GitHub. For compiled, ready-to-use dataframes built from these raw results (and the code that produced them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.100K<n<1M0 likes14k downloads2mo agoHugging Face15BLINK-Benchmark /BLINK BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.image1K<n<10K49 likes14k downloads1y agoHugging Face16clamp-benchmark /clamp-benchmark CLAMP: A Sim-to-Real Benchmark for Closed-Loop Kinematic Pose Estimation and Assembly Reasoning Closed-Loop Assembly and Mechanism Perception Accepted at NeurIPS 2026 (Evaluations & Datasets Track) Kevin Murray1, Randolph Beauregard Robert III2, Petar Z Duric1, Zoran Duric3 1Overlab LLC &nbsp; 2AVA Labs &nbsp; 3George Mason University 📦 Code: https://github.com/overlab-kevin/clamp Your browser does not support the video tag. Overview of all 210 labeled real test scenes.… See the full description on the dataset page: https://huggingface.co/datasets/clamp-benchmark/clamp-benchmark.keypoint-detection1M<n<10M1 likes13k downloads5d agoHugging Face17ydshieh /cache_benchmark0 likes13k downloads2y agoHugging Face18MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K49 likes13k downloads2mo agoHugging Face19DabbyOWL /PDE_Inverse_Problem_Benchmarking PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems. Code: GitHub - ASK-Berkeley/PDEInvBench Sample Usage You can use the provided script from the codebase to batch download the data: pip install huggingface_hub python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.imageother100M<n<1B3 likes12k downloads4mo agoHugging Face20leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes12k downloads6mo agoHugging Face21simular-ai /sai-osworld-v2-benchmark-runs Sai on OSWorld-V2 — benchmark runs of record Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API protocol logs, and run manifests. Run Date Tasks scored Mean score Perfect (1.0) Zeros run1/ 2026-08-12 108/108 0.7276 28 7 run2/ 2026-08-20 108/108 0.7329 33 5 Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.textother2 likes11k downloads1mo agoHugging Face22gaia-benchmark /GAIAgated GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.audion<1K865 likes11k downloads11mo agoHugging Face23clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes11k downloads4y agoHugging Face24koifisharriet /KAIST-Multispectral-Pedestrian-Benchmark1 likes11k downloads2y agoHugging Face25SLM-Lab /benchmark SLM Lab Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book Foundations of Deep Reinforcement Learning. Documentation · Benchmark Results NOTE: v5.0 updates to Gymnasium, uv tooling, and modern dependencies with ARM support - see CHANGELOG.md. Book readers: git checkout v4.1.1 for Foundations of Deep Reinforcement Learning code. BeamRider Breakout KungFuMaster MsPacman Pong Qbert Seaquest Sp.Invaders… See the full description on the dataset page: https://huggingface.co/datasets/SLM-Lab/benchmark.image1K<n<10K0 likes10k downloads7mo agoHugging Face26hffordata /h3-fewstep-benchmark-20260920 H3 few-step benchmark Fifteen few-step checkpoints for the MiniMax-H3 video+audio model, each run on the same 48 English prompts × 2 seeds (0 and 1) × 3 step counts (4, 8, 32) = 288 videos per checkpoint, 4,320 videos in total. Conditioning is text only (no input image, no camera trajectory). Every video is 1344×768, 5 seconds (120 frames at 24 fps), with synthesized audio unless noted. All checkpoints are adapters or distilled variants of MiniMaxAI/MiniMax-H3, except the… See the full description on the dataset page: https://huggingface.co/datasets/hffordata/h3-fewstep-benchmark-20260920.videotext-to-video0 likes10k downloads14d agoHugging Face27UMIbenchmark /UMI-Benchmark-v1 UMI-Benchmark-v1 UMI-Benchmark-v1 contains 20,000 real-world robot manipulation sessions across 10 tasks: 4 single-arm tasks and 6 dual-arm tasks. Dataset Summary Category Tasks Sessions Single-arm 4 9,988 Dual-arm 6 10,012 Total 10 20,000 Tasks ID Folder Type Task Sessions T1 single_arm_task1 Single-arm Stack Baskets 3,000 T2 single_arm_task2 Single-arm Trash Bag 1,991 T3 single_arm_task3 Single-arm Stamp Ink 2… See the full description on the dataset page: https://huggingface.co/datasets/UMIbenchmark/UMI-Benchmark-v1.text10K<n<100K3 likes10k downloads8d agoHugging Face28Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M36 likes9.8k downloads1mo agoHugging Face29iprotogeros /cornetto-benchmarkdataset for Cornetto: A benchmark for LLM-Driven network configuration repair Paper: https://arxiv.org/abs/2604.22513 3 likes9.6k downloads5mo agoHugging Face30clip-benchmark /wds_imagenet-aimage1K<n<10K0 likes9.4k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.