Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tom-jerry-123 /Physical-AI-AV-US PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 2,789,773 samples from 150 000 driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the United States. Format WebDataset — 100 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.png Front-facing wide-angle camera frame (640 × 360 px) {key}.json Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.imagerobotics1M<n<10M0 likes12k downloads7mo agoHugging Face02AVSim /simulation-package Simulation runtime: the third-party half of the AVSim data generation package This repository holds the third-party runtime of the Simulation data generation package, version 3: the pieces the pipeline needs that were not written by the authors. It contains no code and no data of the authors. It is not usable on its own: the private half of the package, AVSim/simulation, downloads this repository into the same directory at a pinned revision with its fetch_runtime.sh script and… See the full description on the dataset page: https://huggingface.co/datasets/AVSim/simulation-package.imagen<1K0 likes3.7k downloads29d agoHugging Face03Avi2006 /spatial-moe-resultsimage10K<n<100K0 likes2.8k downloads8d agoHugging Face04microsoft /AVGen-Bench AVGen-Bench Generated Videos Data Card Overview This data card describes the generated audio-video outputs stored directly in the repository root by model directory. The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.imagetext-to-video1K<n<10K7 likes2k downloads5mo agoHugging Face05Voxel51 /AVM_Segmentation_train Dataset Card for AVM (Around View Monitoring) Semantic Segmentation Dataset This repository provides a FiftyOne-compatible version of the AVM semantic segmentation dataset for autonomous parking systems, with enhanced metadata and visualization capabilities. This is a FiftyOne dataset with 6763 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/AVM_Segmentation_train.imageimage-classification1K<n<10K1 likes1.5k downloads1y agoHugging Face06Ava2000 /Illustrious_Lora_LegacyA directory of all my old Illustrious models (mostly the ones posted on Civitai). They are also hosted here: https://civitai.red/user/Ava_Choco (but censorship is so brutal that it alsmost flags anything anime/manga) Your should be able to dig up the activation tag from the images or from the meta data of the lora. If you can't find it that way, let me know and I will look it up. LoRA Usage Disclaimer This LoRA model is provided as-is for non-commercial use only. Important: The user does not… See the full description on the dataset page: https://huggingface.co/datasets/Ava2000/Illustrious_Lora_Legacy.imagen<1K2 likes1.3k downloads3mo agoHugging Face07Iceclear /AVAAVA: A Large-Scale Database for Aesthetic Visual Analysis See Github Page for tags. Citation @inproceedings{murray2012ava, title={AVA: A large-scale database for aesthetic visual analysis}, author={Murray, Naila and Marchesotti, Luca and Perronnin, Florent}, booktitle={CVPR}, year={2012}, } image14 likes1.1k downloads3y agoHugging Face08Avalon-S /PInVerify PInVerify Dataset An offline embodied benchmark for Active Instance Verification (AIV). Paper arXiv:2605.30639 Code github.com/Avalon-S/PInVerify Project page avalon-s.github.io/PInVerify Venue FMEA Workshop @ CVPR 2026 (Poster) Overview An agent that navigates to a target object does not always arrive at the right instance. Telling "white floral" from "white striped" takes a close look from more than one viewpoint, which is a separate… See the full description on the dataset page: https://huggingface.co/datasets/Avalon-S/PInVerify.imagevisual-question-answering10K<n<100K0 likes1.1k downloads1mo agoHugging Face09aviadcohz /TextureADE TextureADE Real scenes carrying several appearance transitions each, mined from the ADE20K validation split. One of the four evaluation routes in the ICLR 2027 submission on sub-semantic image segmentation: partitioning an image into regions that are coherent in appearance and describable in language, but that need not correspond to any object, part or material class. Images: 212 Code: github.com/aviadcohz/Qwen2SAM_Detecture_Benchmark Weights: aviadcohz/Detecture-ICLR-2027 All… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/TextureADE.imageimage-segmentation1K<n<10K0 likes806 downloads1mo agoHugging Face10initialneil /DREAMS-AVATAR DREAMS-AVATAR The DREAMS-Avatar dataset from the DEGAS paper (3DV 2025), re-registered to pure SMPL-X. These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS (Fig. 1b); what is new here is the registration. 32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted to all views at once by our multiview tracker: 300 shape coefficients, 100 expression coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.imageimage-to-3dn<1K0 likes804 downloads2mo agoHugging Face11Ava2000 /Pony_LoraA directory of all my old Pony models. They are also hosted here: https://civitai.red/user/Ava_Choco (but censorship is so brutal that it alsmost flags anything anime/manga) Your should be able to dig up the activation tag from the images or from the meta data of the lora. If you can't find it that way, let me know and I will look it up. LoRA Usage Disclaimer This LoRA model is provided as-is for non-commercial use only. Important: The user does not claim ownership of the training data used… See the full description on the dataset page: https://huggingface.co/datasets/Ava2000/Pony_Lora.imagen<1K1 likes767 downloads3mo agoHugging Face12Mar-rill /AVDD-TCMI-dataset AVDD-TCMI Dataset Introduction The AVDD-TCMI dataset is a multimodal dataset for depression detection. Our dataset comprises a total of 1,230 valid samples, including 969 non-depressed individuals and 261 depressed individuals. It contains two primary subsets: HDS (Hospital Depression Subset): HDS is collected from volunteers at West China Hospital of Sichuan University and includes 282 valid samples, comprising 159 non-depressed volunteers (including doctors, nurses… See the full description on the dataset page: https://huggingface.co/datasets/Mar-rill/AVDD-TCMI-dataset.image1K<n<10K1 likes747 downloads1y agoHugging Face13trojblue /AVA-Huggingface AVA-Huggingface This repository contains a Hugging Face dataset built from the AVA (Aesthetic Visual Analysis) dataset. The dataset includes images along with their aesthetic scores, total votes, and rating distributions. The data is prepared by filtering out images with fewer than 50 votes and stratifying them based on the computed mean aesthetic score. Dataset Overview Image ID: Unique identifier for each image. Image: The actual image loaded from disk. Mean… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-Huggingface.image100K<n<1M3 likes685 downloads26d agoHugging Face14ahmedtawfik /dfki-av-gopro-hand-objectimage1K<n<10K0 likes624 downloads4mo agoHugging Face15oakmindai /minimax_h3_avatar_500 Watch the full 500-video showcase on YouTube MiniMax H3 Avatar 500 An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset. Version 1.0 · Released August 14, 2026 Dataset contents Each example contains: A 1024 × 1024 reference avatar image A detailed English generation prompt A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.imagen<1K3 likes547 downloads2mo agoHugging Face16trojblue /AVA-aesthetics-10pct-min50-10bins AVA Aesthetics 10% Subset (min50, 10 bins) This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores. Dataset Overview Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins) Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.image10K<n<100K1 likes530 downloads2y agoHugging Face17UnFaZeD07 /AVSBenchimage0 likes497 downloads8mo agoHugging Face18avnishs17 /food_not_food Food vs Not Food Dataset (from Hugging Face ImageNet-1K) This dataset is a binary classification subset derived from the Hugging Face imagenet-1k dataset. It is curated to support the task of distinguishing food images from non-food images. 📦 Dataset Overview Source: imagenet-1k on Hugging Face Datasets Classes: food: 40 selected ImageNet classes representing food items (e.g., pizza, banana, hotdog) not_food: 40 selected classes not related to food (e.g., car, clock… See the full description on the dataset page: https://huggingface.co/datasets/avnishs17/food_not_food.imageimage-classification1K<n<10K1 likes487 downloads1y agoHugging Face19Yuan-avs /Nurisk Nurisk: VQA for Risk Assessment in Autonomous Driving Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains: image: a BEV image question: a driving-related question answer: the ground truth answer Paper NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 . Framework Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Yuan-avs/Nurisk.imagequestion-answering10K<n<100K4 likes471 downloads4mo agoHugging Face20avinashhm /the-welding-defect-dataset-v2image1K<n<10K4 likes465 downloads1y agoHugging Face21tom-jerry-123 /Physical-AI-AV-DE PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 324,105 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 10 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.imagerobotics100K<n<1M0 likes398 downloads6mo agoHugging Face22aviadcohz /RWTD-COCO RWTD-COCO Single natural appearance transitions built from COCO-Stuff by deterministic reuse of human annotation. One of the four evaluation routes in the ICLR 2027 submission on sub-semantic image segmentation: partitioning an image into regions that are coherent in appearance and describable in language, but that need not correspond to any object, part or material class. Images: 256 Code: github.com/aviadcohz/Qwen2SAM_Detecture_Benchmark Weights: aviadcohz/Detecture-ICLR-2027… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/RWTD-COCO.imageimage-segmentation1K<n<10K0 likes395 downloads1mo agoHugging Face23sindhuhegde /avs-spot Dataset Card for AVS-Spot Benchmark This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild" 📝 ArXiv: https://arxiv.org/abs/2503.22668 🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal 💻 Code: https://github.com/Sindhu-Hegde/jegal We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.imagevideo-text-to-textn<1K2 likes367 downloads1y agoHugging Face24AV-Odyssey /AV_Odyssey_Bench_LMMs_Evalaudio1K<n<10K1 likes349 downloads2y agoHugging Face25AV-Odyssey /AV_Odyssey_BenchOfficial dataset for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?". 🌟 For more details, please refer to the project page with data examples: https://av-odyssey.github.io/. [🌐 Webpage] [📖 Paper] [🤗 Huggingface AV-Odyssey Dataset] [🤗 Huggingface Deaftest Dataset] [🏆 Leaderboard] 🔥 News 2024.11.24 🌟 We release AV-Odyssey, the first-ever comprehensive evaluation benchmark to explore whether MLLMs really understand audio-visual… See the full description on the dataset page: https://huggingface.co/datasets/AV-Odyssey/AV_Odyssey_Bench.audioquestion-answeringn<1K5 likes340 downloads2y agoHugging Face26avihayamor /social-instagram-marketing Social — Instagram Marketing Multimodal Dataset A synthetic, multimodal dataset for Social, an AI Instagram-marketing agent. Every row is a single Instagram post idea that pairs a marketing caption with a matching AI-generated image, conditioned on a business brief and brand preferences. Agent pattern: owner brief + brand preferences → 3 similar successful posts (retrieval / recommendation) + 1 freshly generated post (caption + image). Rows (total) 1,447… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/social-instagram-marketing.imagetext-to-image1K<n<10K1 likes339 downloads4mo agoHugging Face27av120 /panda-pick-place-lerobot-14-18-58_01-06-2026 Franka Panda Pick-and-Place — LeRobot v3 Dataset Visuomotor behavior-cloning dataset collected in MuJoCo with a simulated Franka Emika Panda arm. Recorded in LeRobot v3 format (Parquet + MP4 shards). Load from lerobot.datasets import LeRobotDataset ds = LeRobotDataset("av120/panda-pick-place-lerobot-14-18-58_01-06-2026") Task Pick up a cube and place it ~30 cm to the side using a scripted IK state-machine demonstrator. Box position is randomized ±5… See the full description on the dataset page: https://huggingface.co/datasets/av120/panda-pick-place-lerobot-14-18-58_01-06-2026.imageroboticsn<1K0 likes307 downloads4mo agoHugging Face28AvoCahDoe /llava-15-rlmpq-vlm-eval-results RL-MPQ VLM Evaluation Artifacts Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation. Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results Collections (by base VLM) RL-MPQ VLM — LLaVA-1.5-13B — HF collection RL-MPQ VLM — LLaVA-1.5-7B — HF collection RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection RL-MPQ VLM — Qwen2-VL-7B — HF collection Model repos RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.imagevisual-question-answeringn<1K0 likes274 downloads4mo agoHugging Face29ddecosmo /AVA-third-eye-subsetsCategories of Interest Aesthetic Natural 14 Landscape 15 Nature 27 Rural 7 Sky 28 Water Other aesthetic outdoor images 2 Cityscape 61 Street 39 Transportation 10 Urban Technical Of interest 51 Blur 64 Camera Phones 8 Snapshot Other 6 Interior image10K<n<100K0 likes269 downloads6mo agoHugging Face30aviadcohz /Detecture_Benchmarking Detecture Benchmarking Suite A five-dataset benchmark suite for VLM-guided multi-texture segmentation, released alongside the Detecture architecture (VLM-guided multi-texture segmentation via multiplexed grounding). This bundle contains the training set, one in-domain test set, and three out-of-domain evaluation benchmarks used to score Detecture against baseline model families (SAM 3 vanilla, Grounded-SAM 3, Sa2VA, and Qwen2SAM zero-shot) in the Detecture paper. Layout… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/Detecture_Benchmarking.imageimage-segmentation1K<n<10K0 likes258 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.