Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M37 likes9.9k downloads1mo agoHugging Face02langminer /watermark-benchmark-v18 Multi-Class Watermark & Camera Stamp Dataset (Round 18) This repository contains the complete dataset, augmentation assets, and real-world evaluation benchmarks used to train the Champion 3-Class Watermark Classifier (wm_3class_v18_scratch.pt). The dataset addresses a critical problem in media ingestion: automatically excluding produced media, broadcast stills, and stock photography without falsely excluding authentic personal photographs (0.00% false alarms on personal photos… See the full description on the dataset page: https://huggingface.co/datasets/langminer/watermark-benchmark-v18.imageimage-classification10K<n<100K0 likes258 downloads8d agoHugging Face03khadijah00 /ppe-benchmark-eval PPE Benchmark Eval Set (v1) A held-out, human-verified benchmark for evaluating vision-language models on personal protective equipment (PPE) detection — specifically hardhat and safety-vest presence — framed as a VQA-style classification task. What this is 96 images, balanced 24/24/24/24 across the four hardhat × vest combinations (yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of the karabuk-university PPE dataset on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.imagevisual-question-answeringn<1K0 likes188 downloads2mo agoHugging Face04xiapk7 /ANT-A-Benchmark 📊 ANT-A Benchmark This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub. ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample — no forced majority vote — supporting research on soft labels and annotator disagreement. 🗂️ Dataset Structure Each task is provided as two… See the full description on the dataset page: https://huggingface.co/datasets/xiapk7/ANT-A-Benchmark.imagetext-classification10K<n<100K0 likes176 downloads12d agoHugging Face05FForty7 /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,355,161 human responses, collected with the Rapidata Python SDK, comparing how well 30 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.imagetext-to-image100K<n<1M1 likes129 downloads4mo agoHugging Face06kenyag /CoVAtt-Benchmark CoVAtt-Benchmark A large-scale benchmark for generated-image attribution: given an image that was produced by some text-to-image model, decide which model produced it, and decide whether it came from a model the system has never seen before. What this dataset is This is the dataset used to train and evaluate CoVAtt (Content-Based Verification for Attribution of AI-Generated Images, BMVC 2026). CoVAtt is a Siamese network that takes a pair of images and predicts… See the full description on the dataset page: https://huggingface.co/datasets/kenyag/CoVAtt-Benchmark.textimage-classification100K<n<1M0 likes127 downloads2mo agoHugging Face07nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes126 downloads2mo agoHugging Face08thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes123 downloads5mo agoHugging Face09HindsboNikolaj /scope-benchmark SCOPE Benchmark Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories. The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE. When you chain a language model and a vision model together, how do you know which one failed? Contents scope-benchmark/ scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.documentvisual-question-answeringn<1K0 likes112 downloads4mo agoHugging Face10sukiewang /poi-benchmark POI Benchmark: Multi-City Multimodal Points of Interest A large-scale multimodal benchmark pairing Points of Interest (POIs) with street-view imagery, aerial grid photos, and satellite imagery across 10 major cities on 3 continents. Cities Beijing, Chengdu, Guangzhou, Hong Kong, Shanghai, Shenzhen, London, Melbourne, New York, Sydney. Contents Path Size Type Description metadata_aligned.tar 8.6 GB 11 JSON files Enriched & aligned POI metadata per… See the full description on the dataset page: https://huggingface.co/datasets/sukiewang/poi-benchmark.imageimage-to-text100K<n<1M0 likes110 downloads5mo agoHugging Face11ucsahin /Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace. The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels. The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.imageimage-to-text10K<n<100K7 likes103 downloads2y agoHugging Face12Rapidata /Face_Generation_Benchmark Rapidata Human Face Generation Alignment This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating 12 different image generation models on which one can generate faces more accurately. The question that the annotators get asked is: "Which Image follows the description of the human better?" To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.imagetext-to-image1K<n<10K16 likes102 downloads11mo agoHugging Face13nutrientdocs /document-classification-benchmark Document Classification Benchmark (open-vocab, zero-shot) Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot, open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked into a head. Test split only; not for training. Every image is drawn from a permissively-licensed, redistributable source. Powers the document-classification-leaderboard and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.imagezero-shot-image-classification1K<n<10K0 likes97 downloads2mo agoHugging Face14SHPDRG /medical-imaging-model-evaluation-benchmark 医学影像多模态模型评测集(精选示例版) 这是一个面向医学多模态大模型的高质量影像评测集,专门测试模型能否把“看见影像”进一步转化为可解释、可复核、符合临床语境的判断与表达。数据将医学影像与患者描述、病史摘要、检查信息或结构化临床资料配对,覆盖从影像分类、报告生成,到鉴别诊断、治疗方案和胸片质量控制的完整评测链路。 本次公开版本从 2026-07-22 质检通过产物中整理而来,按每个子集最多 50 题进行分层抽样;题量不足 50 的影像质量控制子集完整保留。因此,公开版本包含 5 个任务子集、213 题和 455 个配套影像文件,适合作为医学视觉语言模型的快速对比集、回归测试集和研究教学样例。 数据集亮点 多模态对齐:每条样例同时提供影像和结构化的 question、answer、explanation,支持检查视觉理解、临床语义整合与解释质量。 任务覆盖完整:从“影像是什么”到“如何描述、如何鉴别、如何处置”,并加入真实影像工作流中的胸片质量控制任务。 影像类型丰富:覆盖 X… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/medical-imaging-model-evaluation-benchmark.imageimage-classificationn<1K0 likes86 downloads2d agoHugging Face15DebdipCS /Latent-Resonance-AI-Image-Forensics-Benchmark-N1000 Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000) Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026) 1. Executive Summary & Diagnostic Suite This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.tabularimage-classification1K<n<10K0 likes81 downloads27d agoHugging Face16nutrientdocs /doc-openvocab-benchmark Open-Vocab Document & Figure Classification Benchmark Given a document or figure image and an arbitrary set of text labels, which one is right? This is a zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.imagezero-shot-image-classification1K<n<10K2 likes76 downloads3mo agoHugging Face17DebdipCS /Latent-Resonance-AI-Image-Forensics-Benchmark-N100 Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale) Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026) Benchmark Overview This repository provides: The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.imageimage-classificationn<1K0 likes70 downloads27d agoHugging Face18EDAnonSubmission /benchmark EditJudge-Bench EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as automated judges for image-edit verification. Each row contains a source image, an edited image, a factual edit instruction, counterfactual instructions, and ground-truth scene parameters produced by a controlled Blender/Infinigen generation pipeline. This repository is an anonymous review release for a NeurIPS Evaluations and Datasets submission. Dataset Contents 1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.imageimage-classification1K<n<10K0 likes53 downloads5mo agoHugging Face19Robo531 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes37 downloads7mo agoHugging Face20macular /diabetic-retinopathy-screening-benchmark-africa DR-Africa-Benchmark — Screening-Prevalence-Corrected, Fairness-Instrumented DR Evaluation An evaluation benchmark for diabetic-retinopathy grading under African screening conditions. It does not introduce new labels; it introduces evaluation validity — per-record importance weights that reweight a referral-skewed image set to real Sub-Saharan-Africa population prevalence, plus synthetic subgroup metadata for fairness reporting. Version 1.0.0 · core dr_synth 1.0.0 · part of the… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-screening-benchmark-africa.imageimage-classification1K<n<10K0 likes36 downloads4mo agoHugging Face21BDRC /tibetan-script-classification-benchmark Tibetan Script Classification Benchmark Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training. All images are BDRC manuscript page scans, balanced by subclass. Class Images Subclasses Danyig 60 DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6 Druma 60 Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6 Gyuyig 60 Khyuyig: 31, Tsumachug: 15, Yigchung: 14 Pedri 60 Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.imageimage-classificationn<1K0 likes31 downloads3mo agoHugging Face22oliveirabruno01 /sheep-facial-expression-benchmark Sheep Facial Expression Benchmark Prepared OpenFARM sheep facial-expression benchmark data from the public Mendeley Data record 10.17632/y5sm4smnfr.5. Source Source dataset: https://data.mendeley.com/datasets/y5sm4smnfr Source DOI: 10.17632/y5sm4smnfr.5 Related paper DOI: 10.1016/j.compag.2020.105528 License: CC BY 4.0 Splits { "train": 172, "test": 74, "train_raw": 898, "test_raw": 225 } train and test are filtered/balanced views for benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/sheep-facial-expression-benchmark.imageimage-classification1K<n<10K0 likes23 downloads5mo agoHugging Face23dcher95 /multi-species-benchmark multi-species benchmark Photographs where 2+ species appear in the same frame. Designed to evaluate multi-label species identification and steering capabilities of biological vision-language models. Two sources, unified into one parquet schema. Sources inat21_multilabel (299 rows, 147 images) In-distribution: drawn from iNat21 validation images that already carry an iNat-supplied primary label. We use InternVL3-AWQ to surface images that also… See the full description on the dataset page: https://huggingface.co/datasets/dcher95/multi-species-benchmark.imageimage-classification1K<n<10K0 likes21 downloads4mo agoHugging Face24ash12321 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes19 downloads9mo agoHugging Face25anonymous-vision-bench /vision-benchmarkgated GravCal: Large-Scale Orientation-Diverse Dataset for IMU Gravity Calibration NeurIPS 2026 Evaluations & Datasets Track Dataset Description GravCal is a large-scale dataset specifically designed for single-image IMU gravity calibration. The dataset addresses a critical gap in existing visual-inertial datasets, which exhibit severe upright-pose bias with most frames captured near canonical orientations. Key Features 148,000+ frames with diverse camera… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-vision-bench/vision-benchmark.textimage-classificationn<1K0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.