Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K207 likes2.1m downloads7h agoHugging Face02banned-historical-archives /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.imagen<1K93 likes1.3m downloads1y agoHugging Face03BuLei /imgbedimagen<1K0 likes940k downloads7d agoHugging Face04xlangai /osworld_v2_assetsimagen<1K17 likes592k downloads2mo agoHugging Face05updatebao /geonamebase_1image26 likes584k downloads1y agoHugging Face06HuggingFaceM4 /the_cauldron Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. Load the dataset To load the dataset, install the library datasets with pip install datasets. Then, from datasets import load_dataset ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d") to download… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/the_cauldron.image1M<n<10M561 likes395k downloads2y agoHugging Face07agents-course /course-imagesimagen<1K21 likes333k downloads1y agoHugging Face08huggingface-course /documentation-imagesimagen<1K3 likes317k downloads1y agoHugging Face09jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes292k downloads3y agoHugging Face10fujinchu /imgbedimagen<1K1 likes235k downloads4h agoHugging Face11huggingface /badges Badges A set of badges you can use anywhere. Just update the anchor URL to point to the correct action for your Space. Light or dark background with 4 sizes available: small, medium, large, and extra large. How to use? With markdown, just copy the badge from: https://huggingface.co/datasets/huggingface/badges/blob/main/README.md?code=true With HTML, inspect this page with your web browser and copy the outer html. Available sizes Small Medium… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/badges.imagen<1K50 likes212k downloads8mo agoHugging Face12RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M23 likes205k downloads23d agoHugging Face13harborframework /terminal-bench-2.1 Terminal-Bench 2.1 (Harbor git-repos dataset) Harbor website · Harbor GitHub This is a private mirror of the task content from harbor-framework/terminal-bench-2-1 at commit 7131e43 (the source repo has no tagged releases yet), laid out so it can be consumed directly by Harbor's git-repos dataset support. The primary source is the GitHub repository above — please open issues and pull requests there, not here. How to run Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.documentn<1K15 likes188k downloads21d agoHugging Face14GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K46 likes185k downloads47m agoHugging Face15uoft-cs /cifar10 Dataset Card for CIFAR-10 Dataset Summary The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.imageimage-classification10K<n<100K126 likes177k downloads3y agoHugging Face16mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K41 likes165k downloads1mo agoHugging Face17wyu1 /Leopard-Instruct Leopard-Instruct Paper | Github | Models-LLaVA | Models-Idefics2 Summaries Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint]. Loading dataset to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.image1M<n<10M64 likes162k downloads2y agoHugging Face18huggingface /brand-assetsimagen<1K13 likes149k downloads7mo agoHugging Face19Narsil /image_dummy\audion<1K0 likes148k downloads5y agoHugging Face20worldbenchmark /IntuitivePhysics WorldBench WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models and vision-language models. There are two components: Video based: This is a benchmark designed to evaluate video-to-video world foundation models such as Cosmos. It consists of videos 132 frames long of 425 simulated scenes with RGB, Normals, Depth, Flow, and Segmentations. Text based: This is a subset which adds text-based questions to 181 videos from… See the full description on the dataset page: https://huggingface.co/datasets/worldbenchmark/IntuitivePhysics.imagevisual-question-answering100K<n<1M2 likes139k downloads5mo agoHugging Face21harborframework /terminal-bench Terminal-Bench The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published version; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.documentn<1K5 likes130k downloads21d agoHugging Face22hf-vision /course-assetsimagen<1K9 likes130k downloads2y agoHugging Face23ESA-philab /OceanDepths OceanDepths GeoTIFF Raster and Aligned ARGO Dataset This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.image1M<n<10M0 likes128k downloads2mo agoHugging Face24jamesqijingsong /chengyu時間: 2018年做成網站 https://chengyu.18dao.net 2024年用AI將文本生成圖片 2025年上傳到Hugging Face的Datasets 数据集中的文件总数: 20609 目录 "Text-to-Image/" 下的文件数量: 10296,子目錄數:5148,每個子目錄兩個文件,一個原始的文生圖png圖片,一個圖片解釋txt文件 目录 "image-chengyu/" 下的文件数量: 5155,加字的圖片jpg文件 目录 "text-chengyu/" 下的文件数量: 5156,文字解釋txt文件 image1K<n<10K1 likes124k downloads2y agoHugging Face25andaba /RT-PosePaper RT-Pose: A 4D Radar Tensor-based 3D Human Pose Estimation and Localization Benchmark (ECCV 2024) RT-Pose introduces a human pose estimation (HPE) dataset and benchmark by integrating a unique combination of calibrated radar ADC data, 4D radar tensors, stereo RGB images, and LiDAR point clouds. This integration marks a significant advancement in studying human pose analysis through multi-modality datasets. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/andaba/RT-Pose.imagekeypoint-detection100K<n<1M22 likes123k downloads2y agoHugging Face26HorizonRobotics /EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer 3d1K<n<10K7 likes120k downloads3mo agoHugging Face27jamesqijingsong /zidian时间线: 2018年搭建成网站 https://zidian.18dao.net 2024年使用AI技術為《國語字典》生成配圖。 2025年上傳到Hugging Face做成數據集。 数据集中的文件: 目录 "image/" 下的文件数量: 4307,文生圖原始png圖片 目录 "image-zidian/" 下的文件数量: 4307,加字後的jpg圖片 目录 "text-zidian/" 下的文件数量: 4307,圖片解釋文字 目录 "pinyin/" 下的文件数量: 1702,拼音mp3文件 audio1K<n<10K0 likes118k downloads2y agoHugging Face28dalle-mini /witimage1M<n<10M7 likes116k downloads5y agoHugging Face29ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes113k downloads1y agoHugging Face30ShadenA /MathNet Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1. Quick start from datasets import load_dataset # Default: all problems ds = load_dataset("ShadenA/MathNet", split="train") # Or a specific country / competition-body config… See the full description on the dataset page: https://huggingface.co/datasets/ShadenA/MathNet.imagequestion-answering10K<n<100K96 likes111k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.