Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kaphathy /Dataset MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records 1. Executive Summary & Repository Overview The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.textimage-classificationn<1K2 likes2.1k downloads12d agoHugging Face02ken-jo /qev-data QEV data Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model. Model. Configurations overlap. Do not concatenate them or assume independent test sets. ZIPs in corpora/ use qev-<stage>.zip and contain eligible original records, images and license notices. Extracted stage folders and original record/source IDs preserve the recorded training provenance. Viewer rows expose request/target schemas as JSON strings; parse with json.loads. IDs, group IDs… See the full description on the dataset page: https://huggingface.co/datasets/ken-jo/qev-data.imagetext-classification10K<n<100K0 likes1.1k downloads9d agoHugging Face03paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face04paulpacaud /rlbenchfail_train_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.tabularvisual-question-answering10K<n<100K0 likes800 downloads8mo agoHugging Face05paulpacaud /rlbenchfail_val_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.tabularvisual-question-answering1K<n<10K0 likes547 downloads8mo agoHugging Face06Davidup1 /GeoChrono-Data ChronoBench & ChronoInstruct ChronoBench is a comprehensive, multi-dimensional, and multi-granularity benchmark for high-resolution long-temporal remote sensing understanding. It decomposes long-term remote sensing understanding into a four-level cognitive hierarchy — from Land Cover Perception through Temporal Recognition and Long-Term Memory to Spatio-Temporal Reasoning — comprising 12 sub-tasks and 17,689 rigorously validated QA pairs derived from 3,469 high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Davidup1/GeoChrono-Data.tabularvisual-question-answering10K<n<100K0 likes295 downloads2mo agoHugging Face07Harisundar /PALL-VLM-data PALL-VLM-data — Dental Vision-Language Dataset The training dataset for Harisundar/PALL-VLM, a dental vision-language model. It contains 32,884 records over 52,461 images, formatted as image+text conversations for LLaVA-style instruction tuning. Curated by: Harisundar R Used by: Harisundar/PALL-VLM · PALL on GitHub Language: English Layout vlm_train/ ├── images/ # 52,461 dental images ├── train.jsonl # 29,667 records ├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.imageimage-text-to-text10K<n<100K1 likes208 downloads4mo agoHugging Face08paulpacaud /ur5fail_test_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.tabularvisual-question-answeringn<1K1 likes203 downloads8mo agoHugging Face09paulpacaud /Guardian-FailCoT-OOD-datasets Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026): UR5-Fail — our newly collected three-view real-robot benchmark. RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023). RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.tabularvisual-question-answering1K<n<10K1 likes198 downloads5mo agoHugging Face10paulpacaud /ur5fail_train_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.tabularvisual-question-answering1K<n<10K0 likes190 downloads8mo agoHugging Face11paulpacaud /ur5fail_val_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_val_dataset.tabularvisual-question-answeringn<1K0 likes174 downloads8mo agoHugging Face12paulpacaud /bdv2fail_train_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.tabularvisual-question-answering10K<n<100K1 likes141 downloads8mo agoHugging Face13paulpacaud /bdv2fail_val_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_val_dataset.tabularvisual-question-answering1K<n<10K0 likes110 downloads8mo agoHugging Face14paulpacaud /bdv2fail_test_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes104 downloads8mo agoHugging Face15Saransh-cpp /orion-dataset Dataset The ORION dataset is a curated collection of satellite imagery and triage labels used to fine-tune the VLM for orbital image classification. Images are fetched from SimSat's Mapbox API and paired with classification prompts and ground-truth labels. Dataset Structure images/ low_ocean_pacific_nemo.png med_city_chicago.png high_port_rotterdam.png ... train_dataset.jsonl val_dataset.jsonl test_dataset.jsonl images/: 512x512 RGB satellite images fetched from… See the full description on the dataset page: https://huggingface.co/datasets/Saransh-cpp/orion-dataset.imageimage-classificationn<1K0 likes79 downloads5mo agoHugging Face16inaf-oact-ai /solar-flare-hmi-datasplitsThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of SDO/HMI magnetogram images compiled by Boucheron, L.E., et al., 2023, Sci Data 10, 825, https://doi.org/10.1038/s41597-023-02628-8. Splits "train", "val", "test" corresponds to the original data splits provided by Boucheron et al., while the other splits are created by downsampling the No-Flare and C flare class to obtain more balanced splits and… See the full description on the dataset page: https://huggingface.co/datasets/inaf-oact-ai/solar-flare-hmi-datasplits.textimage-classification1M<n<10M0 likes54 downloads1y agoHugging Face17newazhala /ui-design-audit-dataset UI Design Audit Screenshot Benchmark v2.1 A reproducible synthetic benchmark of 3,000 mobile and web UI screenshots labeled across 12 design-risk categories. Splits train: 2,400 validation: 300 test: 300 Labels small_touch_targets low_contrast action_overload navigation_overload form_friction content_density responsive_risk modal_overuse deep_scrolling weak_hierarchy interaction_overload mobile_web_mismatch Data creation Every… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/ui-design-audit-dataset.imageimage-classification1K<n<10K0 likes52 downloads15d agoHugging Face18Mobiusi /Asparagus-Identification-Dataset Asparagus Identification Dataset The current agricultural industry faces challenges in efficient crop monitoring and recognition. Traditional manual detection methods are inefficient and prone to errors. Existing solutions often rely on empirical judgment without scientific model support. This dataset aims to provide diverse asparagus images to help train automatic recognition models, improving the accuracy and efficiency of crop monitoring. The dataset includes images of asparagus… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Asparagus-Identification-Dataset.textimage-classificationn<1K0 likes38 downloads7mo agoHugging Face19weitaikang /Weitaikang_bench_data HeroFrame-Bench A benchmark for key-frame selection that scores a selected frame directly, without routing it through a question-answering model and without requiring it to match a fixed reference set. 204 films, 2,031 frozen comparison chains, 1,970 learned criteria, 30,465 pairwise judgements. What problem this addresses A film is almost always encountered first as a single still: a cover, a thumbnail, a poster. Producing that still from the film is the task… See the full description on the dataset page: https://huggingface.co/datasets/weitaikang/Weitaikang_bench_data.imageimage-classificationn<1K0 likes34 downloads2mo agoHugging Face20Mobiusi /Pumpkin-Maturity-Grading-Dataset Pumpkin Maturity Grading Dataset The current agricultural industry faces challenges in quality control of crops, especially in judging the maturity of pumpkins. Traditional methods often rely on manual identification, which is inefficient and prone to errors. Although existing image recognition technology has made some progress, there is a lack of high-quality datasets specifically targeted at pumpkin maturity. This dataset aims to improve the accuracy of maturity assessment for… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pumpkin-Maturity-Grading-Dataset.textimage-classificationn<1K0 likes25 downloads7mo agoHugging Face21rlogh /lanternfly_research_dataset Lantern Fly Research Dataset This dataset contains human-verified spotted lanternfly sightings collected through the Lantern Fly Tracker app. Each entry includes high-quality photos, precise geolocation data, and comprehensive metadata for ecological research. 🎯 Purpose This dataset supports: Ecological research on spotted lanternfly distribution and spread patterns Machine learning model training with verified, high-quality data Temporal and spatial analysis of… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_research_dataset.textimage-classificationn<1K0 likes20 downloads1y agoHugging Face22Mobiusi /Pesticide-Spraying-Scene-Classification-Dataset Pesticide Spraying Scene Classification Dataset The current agricultural industry faces challenges such as low spraying efficiency and environmental pollution, especially during large-scale farmland spraying. Traditional methods rely on manual operations, which can lead to pesticide waste and uneven spraying. Existing solutions often lack efficient image recognition technology and cannot monitor spraying effectiveness and crop conditions in real-time. This dataset aims to support… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pesticide-Spraying-Scene-Classification-Dataset.textimage-classificationn<1K0 likes16 downloads7mo agoHugging Face23Mobiusi /Papaya-Tree-Recognition-Dataset Papaya Tree Recognition Dataset The current agricultural sector faces issues of low efficiency in crop recognition and management, especially against the backdrop of the growing development of smart agriculture. Traditional manual recognition methods are unable to meet the rapidly changing needs. Existing solutions often rely on image data from a single environment, lacking diversity and universality, which leads to poor model generalization. The Papaya Tree Recognition Dataset aims… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Papaya-Tree-Recognition-Dataset.textimage-classificationn<1K0 likes13 downloads7mo agoHugging Face24Mobiusi /Raspberry-Variety-Classification-Dataset Raspberry Variety Classification Dataset The current agricultural industry faces challenges in managing the diversity of crop types, especially in raspberry cultivation and variety identification. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to address the issue of low accuracy in variety classification by providing high-quality raspberry images, meeting the needs of intelligent agriculture. The dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Raspberry-Variety-Classification-Dataset.textimage-classificationn<1K0 likes13 downloads7mo agoHugging Face25lcuifer0 /ourdream-database OurDream Character Database Complete database of 20,884 AI characters from ourdream.ai, categorized as Women (12,492) and Trans (8,392). Dataset Structure database/all_characters.json — Full character database with 20,884 entries Each entry includes: id, displayId, name, gender, style, age, likeCount, messageCount, tags, shortDescription, thumbUrl, category Categories Category Count Women 12,492 Trans 8,392 Total 20,884… See the full description on the dataset page: https://huggingface.co/datasets/lcuifer0/ourdream-database.tabulartext-generation10K<n<100K0 likes13 downloads4mo agoHugging Face26shangzx /Pumpkin-Maturity-Grading-Dataset Pumpkin Maturity Grading Dataset The current agricultural industry faces challenges in quality control of crops, especially in judging the maturity of pumpkins. Traditional methods often rely on manual identification, which is inefficient and prone to errors. Although existing image recognition technology has made some progress, there is a lack of high-quality datasets specifically targeted at pumpkin maturity. This dataset aims to improve the accuracy of maturity assessment for… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Pumpkin-Maturity-Grading-Dataset.textimage-classificationn<1K0 likes11 downloads7mo agoHugging Face27Mobiusi /Experiment-Design-Sketch-Image-Classification-Dataset Experiment Design Sketch Image Classification Dataset In the field of industrial manufacturing, the design process often relies on a large number of design sketches that need to be quickly converted into actual engineering designs during the subsequent manufacturing stages. However, manually processing these sketches is often time-consuming and prone to errors, currently relying mainly on manual labeling and conversion by designers, which is inefficient and unstable. Existing… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Experiment-Design-Sketch-Image-Classification-Dataset.textimage-classificationn<1K0 likes11 downloads7mo agoHugging Face28Mobiusi /Field-Pumpkin-Recognition-Dataset Field Pumpkin Recognition Dataset The current agriculture industry faces challenges such as low efficiency in crop recognition and difficulty in pest monitoring. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to support the training of machine learning models by providing high-quality pumpkin image data, enhancing the accuracy and speed of pumpkin recognition. Data is primarily collected in the field using… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Field-Pumpkin-Recognition-Dataset.textimage-classificationn<1K0 likes10 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.