Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tom-jerry-123 /Physical-AI-AV-US PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 2,789,773 samples from 150 000 driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the United States. Format WebDataset — 100 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.png Front-facing wide-angle camera frame (640 × 360 px) {key}.json Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.imagerobotics1M<n<10M0 likes12k downloads7mo agoHugging Face02General-Medical-AI /Ophora-160Kgated Introduction Ophora-160K contains 162,185 video clip-instruction pairs extracted from 9,819 narrative videos of ophthalmic surgery. The average duration of all clips is 5.54 seconds. All ophthalmic narrative videos were collected from the YouTube platform. The file ophora161k.csv contains all video clip IDs and their generation instructions. The file ophora28k.csv contains content that has been filtered to remove sensitive information such as subtitles and watermarks.… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/Ophora-160K.texttext-to-video100K<n<1M1 likes4.5k downloads1y agoHugging Face03ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.5k downloads7mo agoHugging Face04HiDream-ai /ReCo-Data ReCo-Data Dataset Card Introduction ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format. 1. Dataset Statistics Statistics Figure Caption: (a) Overview of scale (b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.textimage-to-video1M<n<10M97 likes2.2k downloads5mo agoHugging Face05laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M16 likes2k downloads11mo agoHugging Face06ai-music /ai-music-deduplicated AI Music Deduplicated A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar. Overview Subset Songs Tar Files Size Audio Format Source Platform mureka ~312K 49 ~981 GB .mp3 Mureka riffusion ~105K 14 ~266 GB .m4a Riffusion sonauto ~15K 2 ~25 GB .ogg Sonauto suno ~307K 65 ~1.3 TB .mp3 Suno udio~126K 33 ~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.audioaudio-classification100K<n<1M5 likes1.6k downloads8mo agoHugging Face07clip-benchmark /wds_fgvc_aircraftimage1K<n<10K0 likes1.1k downloads4y agoHugging Face08backups /ai-maudio10M<n<100M1 likes861 downloads1y agoHugging Face09AIPeanutman /OpenSubject OpenSubject Dataset OpenSubject is a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation, as presented in the paper OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation. Project Page & Code See the main repository for more details and code: OpenSubject Dataset Structure OpenSubject/ ├── Images_packages/ # Compressed image… See the full description on the dataset page: https://huggingface.co/datasets/AIPeanutman/OpenSubject.imageimage-to-image1M<n<10M1 likes843 downloads10mo agoHugging Face10Stemson-AI /Warwick-STEM Warwick STEM Dataset (WebDataset) A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018. Dataset Description This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.imageimage-to-image10K<n<100K2 likes710 downloads6mo agoHugging Face11sleeping-ai /MemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata. We share these files as-part of research initiative. audio100K<n<1M0 likes630 downloads1y agoHugging Face12smz8599 /GUI-AIMA-multiturnimage100K<n<1M1 likes619 downloads8mo agoHugging Face13pumb-ai /synthetic-cyrillic-largeimageimage-to-text1M<n<10M3 likes566 downloads2y agoHugging Face14projecte-aina /parlament_parla_v3 Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.audioautomatic-speech-recognition100K<n<1M1 likes425 downloads2y agoHugging Face15tom-jerry-123 /Physical-AI-AV-DE PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 324,105 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 10 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.imagerobotics100K<n<1M0 likes398 downloads6mo agoHugging Face16zzha6204 /RU-AI-noise RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection This is the noise agumented data for paper: RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection The original dataset is avaliable at zenodo: https://zenodo.org/records/11406538 The official repo is avaliable at: https://github.com/ZhihaoZhang97/RU-AI Reference We are appreciated the open-source community for the datasets and the models. Microsoft COCO: Common Objects in… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/RU-AI-noise.audioany-to-any1M<n<10M2 likes355 downloads2y agoHugging Face17orion-ai-lab /Thalia Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring Paper | GitHub | Interactive Demo (Colab) Thalia is a global, multi-modal dataset for volcanic activity monitoring through Satellite-based Interferometric Synthetic Aperture Radar (InSAR) imagery. Building upon the Hephaestus dataset, Thalia provides higher-resolution, multi-source, and multi-temporal data in a machine-learning-ready format. Dataset Overview Thalia consists of 38 spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Thalia.textimage-classification10K<n<100K3 likes297 downloads5mo agoHugging Face18aiintelligentsystems /vel_commons_wikidata Visual Entity Linking: Wikimedia Commons & Wikidata This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict. Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/ uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.image100K<n<1M5 likes283 downloads2y agoHugging Face19tom-jerry-123 /Physical-AI-AV-FR PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,909 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.imagerobotics10K<n<100K0 likes209 downloads6mo agoHugging Face20tom-jerry-123 /Physical-AI-AV-ES PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,674 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.imagerobotics10K<n<100K0 likes202 downloads6mo agoHugging Face21tom-jerry-123 /Physical-AI-AV-IT PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,991 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.imagerobotics10K<n<100K0 likes200 downloads6mo agoHugging Face22xtcpete /air_groundtext10K<n<100K0 likes161 downloads8mo agoHugging Face23AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes158 downloads3y agoHugging Face24SMIIP-lab /AISHELL6-Whispergated 🗣️ AISHELL6-Whisper AISHELL6-Whisper is a large-scale open-source Chinese Mandarin audio-visual whisper speech dataset,containing 30 hours each of whisper and parallel normal speech, with synchronized frontal RGB facial videos. 📘 Dataset Summary Property Description Language Chinese (Mandarin, ZH) License CC BY-NC-SA 4.0 Duration ~60 hours total (30 h whisper + 30 h normal) Speakers 167 total (121 with RGB-D, 46 audio-only) Environment Controlled… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL6-Whisper.audio10K<n<100K11 likes149 downloads9mo agoHugging Face25AI4Protein /VenusREMtextn<1K0 likes130 downloads2y agoHugging Face268bits-ai /ZOD-Mini-2D-Road-Scenes ZOD-Mini-2D-Road-Scenes The ZOD-Mini-2D-Road-Scenes dataset is derived from the Zenseact Open Dataset (ZOD), property of Zenseact AB (© 2022 Zenseact AB), and is licensed under the permissive CC BY-SA 4.0. Any public use, distribution, or display of this dataset must contain this entire notice: For this dataset, Zenseact AB has taken all reasonable measures to remove all personally identifiable information, including faces and license plates. To the extent that you like to request… See the full description on the dataset page: https://huggingface.co/datasets/8bits-ai/ZOD-Mini-2D-Road-Scenes.image100K<n<1M0 likes119 downloads2y agoHugging Face27mjhbest /aihub-medicine-OCRtextn<1K0 likes117 downloads2y agoHugging Face28bitmind /aislop-videostext1K<n<10K0 likes115 downloads1y agoHugging Face29Ailovejinx /planartrackplustextn<1K0 likes110 downloads2y agoHugging Face30AiArtLab /1024 MidJourney NijiJourney E-Shushu Compressed by a factor of 64 pixels in JPEG, 97 quality, maximum side length of 1024. Mixed labeling using different models: Human prompts in MJ/NJ Long captions (LLaVA) Short captions (LLaVA + LLaMA) Medium captions (Moondream) image100K<n<1M2 likes104 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.