Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cvml-nus /assembly101gated Assembly101 Assembly101 is a procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/cvml-nus/assembly101.textn<1K20 likes52k downloads4mo agoHugging Face02cvssp /WavCaps WavCaps WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset). Paper: https://arxiv.org/abs/2303.17395 Github: https://github.com/XinhaoMei/WavCaps Statistics Data Source # audio avg. audio duration (s)avg. text length FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.textn<1K56 likes47k downloads3y agoHugging Face03junma /CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation. Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage) Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage) CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.3dimage-segmentation24 likes18k downloads7mo agoHugging Face04cvlab /new-york-smells New York Smells: A Large Multimodal Dataset for Olfaction While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory data collected in natural settings. We present New York Smells, a large-scale dataset of paired image and olfactory signals captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.image10K<n<100K1 likes12k downloads2mo agoHugging Face05EPFL-CVLAB-SPACECRAFT /PocketQubeimage0 likes9.9k downloads8mo agoHugging Face06heitorrosa /cvm-corpus CVM Filings Corpus (PT-BR) Brazilian CVM regulatory filings in Portuguese, cleaned and chunked for language-model pretraining. Built for DAPT on financial Portuguese. Contents Path What output/corpus.jsonl Full corpus (7.2 GB). Chunk schema: text, company, cnpj, category, subject, date, year, document_id, chunk_id, extraction_quality. output/corpus-250M.jsonl DSIR-selected 250M-token subset (token count by chars÷4 proxy ≈ 150M whitespace tokens): 187… See the full description on the dataset page: https://huggingface.co/datasets/heitorrosa/cvm-corpus.documenttext-generation100K<n<1M0 likes9.6k downloads8d agoHugging Face07nyu-visionx /CV-Bench Cambrian Vision-Centric Benchmark (CV-Bench) This repository contains the Cambrian Vision-Centric Benchmark (CV-Bench), introduced in Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. Files The test*.parquet files contain the dataset annotations and images pre-loaded for processing with HF Datasets. These can be loaded in 3 different configurations using… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/CV-Bench.imagevisual-question-answering1K<n<10K48 likes7.7k downloads1y agoHugging Face08EPFL-CVLAB-SPACECRAFT /SwissCubeimage10K<n<100K3 likes6.5k downloads2y agoHugging Face09mort666 /cv_corpus_v22 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. NOTE: currently converting to parquet for convenience.. WIP Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.audioautomatic-speech-recognition1M<n<10M0 likes4.9k downloads10mo agoHugging Face10nvidia /cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset. Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes. This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub. textn<1K40 likes4.5k downloads3mo agoHugging Face11mypThu /SportsSlomo-CVS 🎥 SportsSloMo-CVS Dataset This repository contains the dataset presented in the paper Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor. The Complementary Vision Sensor (CVS), known as Tianmouc, captures synchronized RGB frames together with high-frame-rate, multi-bit spatial difference (SD, encoding structural edges) and temporal difference (TD, encoding motion cues) data within a single RGB exposure. This dataset facilitates research in RGB… See the full description on the dataset page: https://huggingface.co/datasets/mypThu/SportsSlomo-CVS.textimage-to-imagen<1K4 likes4.2k downloads4mo agoHugging Face12Beijing-AISI /C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models". We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.texttext-generation100K<n<1M3 likes3.1k downloads1y agoHugging Face13zzsi /cvlimagen<1K0 likes3.1k downloads2d agoHugging Face14CVC2233 /Long-Horizon-GUI-Datasetimage1K<n<10K2 likes3k downloads10mo agoHugging Face15UARK-NED3 /BoilingBench-CV BoilingBench-CV Dataset Version: v0.1.0 Maintainer: NED3 Laboratory, University of Arkansas License: CC BY 4.0 DOI: 10.5281/zenodo.22264378 Mirror of the Zenodo deposit of 3 September 2026, published here because most users of these data work in the Hugging Face ecosystem. The file set was verified identical to the deposit at upload time: 7,147 files, 4.20 GB uncompressed. Authors Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University), Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.imageimage-segmentationn<1K1 likes2.5k downloads20d agoHugging Face16AILab-CVC /SEED-Bench SEED-Bench Card Benchmark details Benchmark type: SEED-Bench is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 19K multiple choice questions with accurate human annotations, which covers 12 evaluation dimensions including the comprehension of both the image and video modality. Benchmark date: SEED-Bench was collected in July 2023. Paper or resources for more information: https://github.com/AILab-CVC/SEED-Bench License:… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Bench.visual-question-answering10K<n<100K24 likes2.1k downloads2y agoHugging Face17VLR-CVC /ComicsPAP Comics: Pick-A-Panel Updated val and test on 25/02/2025 This is the dataset for the ICDAR 2025 Competition on Comics Understanding in the Era of Foundational Models. Please, check out our 🚀 arxiv paper 🚀 for more information 😊 The competition is hosted in the Robust Reading Competition website and the leaderboard is available here. The dataset contains five subtask or skills: Sequence Filling Given a sequence of comic panels, a missing panel, and a set of option panels, the… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/ComicsPAP.image10K<n<100K15 likes1.9k downloads1y agoHugging Face18afaji /cvqa About CVQA CVQA is a culturally diverse multilingual VQA benchmark consisting of over 10,000 questions from 39 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories. This data is designed for use as a test set. Please submit your submission here to evaluate your model performance. CVQA is constructed through a collaborative effort led by a team of researchers from MBZUAI. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/afaji/cvqa.imagequestion-answering10K<n<100K40 likes1.9k downloads2y agoHugging Face19changelinglab /cv-v1.0-segment CommonVoice v1 Phone-Segment Alignments Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the charsiu/cv_ali release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09). Dataset summary lang train rows train hrs val rows val hrs test rows test hrs en 1,008,669 1,354.0 3,537 4.9 1,285 1.7 rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.audioautomatic-speech-recognition1M<n<10M3 likes1.9k downloads6mo agoHugging Face20hitoshura25 /cvefixes CVEfixes Security Vulnerabilities Dataset Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories. Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code. Usage from datasets import load_dataset dataset = load_dataset("hitoshura25/cvefixes") Citation If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/cvefixes.tabulartext-generation10K<n<100K5 likes1.8k downloads1y agoHugging Face21berkaytrhn /cvc-clinicdb CVC-ClinicDB (file mirror) 612 frames extracted from 29 colonoscopy sequences, each with a pixel-level polyp mask (Bernal et al., 2015). Native resolution 384x288. This is a plain file mirror, not a datasets-format repo: images and masks are stored as files so any path-based dataloader can consume them directly after snapshot_download. The dataset viewer is disabled for that reason. Layout CVC-ClinicDB/ Original/*.png # 612 RGB frames, 384x288 Ground… See the full description on the dataset page: https://huggingface.co/datasets/berkaytrhn/cvc-clinicdb.image-segmentationn<1K0 likes1.7k downloads1mo agoHugging Face22StonyBrook-CVLab /doc3D-datasetgated doc3D Doc3D is the first 3D dataset focused on document unwarping with realistic paper warping and renderings. It contains 100k images with the following ground-truths: 3D Coordinates Depth UV Backward Mapping Albedo Normals Checkerboard Useful links: More details of the data usage instructions are available in the GitHub repo: https://github.com/cvlab-stonybrook/doc3D-dataset Link to the training code: https://github.com/cvlab-stonybrook/DewarpNet Link to the… See the full description on the dataset page: https://huggingface.co/datasets/StonyBrook-CVLab/doc3D-dataset.imageimage-to-image9 likes1.7k downloads10mo agoHugging Face23Voxel51 /CVPR_2024_Papers Dataset Card for cvpr2024_papers This is a FiftyOne dataset with 2379 samples. The dataset consists of images of the first page for accepted papers to CVPR 2024, plus their abstract and other metadata. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'split', 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/CVPR_2024_Papers.image1K<n<10K1 likes1.7k downloads2y agoHugging Face24AILab-CVC /SEED-Data-Edit-Part2-3 SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part2-3.text-to-image1M<n<10M12 likes1.7k downloads2y agoHugging Face25CVML-TueAI /Breakfast-Actions 🍳 Breakfast Actions Dataset (HF + WebDataset Ready) This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides: 4 evaluation splits (s1, s2, s3, s4) JSONL metadata describing each video, participant, camera, and frame-level action segments Raw AVI videos stored directly on HuggingFace Optional WebDataset shards for streaming training 📁 Folder Layout Breakfast-Actions/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.text1K<n<10K0 likes1.7k downloads10mo agoHugging Face26anishanish383 /cv-project0 likes1.6k downloads6mo agoHugging Face27pm-science /cv4cdd_4d Content This repository stores the contents of the data/ directory from the following GitLab repository:https://gitlab.uni-mannheim.de/processanalytics/cv4cdd The data is organized as follows: input_cdlgContains training, validation, and test datasets used to train the computer vision models. input_cdriftContains external datasets used to evaluate the trained models. model_training_loggingContains model checkpoints for all training runs. This includes both relevant checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/pm-science/cv4cdd_4d.0 likes1.4k downloads8mo agoHugging Face28SpeechAntiSpoofingBenchmarks /CVoiceFake_small CVoiceFake (small) Benchmark-ready packaging of the CVoiceFake (small) multilingual audio-deepfake detection set introduced with SafeEar (arXiv 2409.09272), for speech anti-spoofing / synthetic-voice detection. Overview CVoiceFake is a multilingual deepfake-audio dataset built on top of Mozilla CommonVoice: genuine clips are re-synthesised by a bank of vocoders to produce matched spoof audio. This repo packages the public small subset — a random ~10% of the full… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/CVoiceFake_small.audioaudio-classification100K<n<1M0 likes1.4k downloads3mo agoHugging Face29EvanOLeary /shinka-cvdp-benchmark-full0 likes1.4k downloads4mo agoHugging Face30VLR-CVC /DocVQA-2026 DocVQA 2026 | ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains Building upon previous DocVQA benchmarks, this evaluation dataset introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. By expanding coverage to new document domains and… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/DocVQA-2026.imagevisual-question-answeringn<1K74 likes1.3k downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.