Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pixelprose /pixelprose-shards PixelProse Sharding Tars arXiv | public-released version: pixelprose | JSON-only version: pixelprose-jsons summary Each tar file is approximately 500-600 MB, friendly for fast on-the-fly sampling, filtering, and loading in dataloaders. Each tar file contains triplets of images, text, and JSON files. The *.txt files contain the raw original captions, while the *.json files include all the relevant information. Due to Gemini-1.0 internal version changes during the… See the full description on the dataset page: https://huggingface.co/datasets/pixelprose/pixelprose-shards.image1M<n<10M2 likes13k downloads10mo agoHugging Face02LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes2.7k downloads6mo agoHugging Face03staturecrane /pixelprose_webp_512image0 likes1.5k downloads2y agoHugging Face04tomg-group-umd /pixelprose From Pixels to Prose: A Large Dataset of Dense Image Captions [ arXiv paper ] | [ 🌮 image tars ] PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions, leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions. 1. Details Total number of image-caption pairs: 16,896,214 (16.9M) 6,538,898 (6.5M) pairs in the split of CommonPool 9,066,455 (9.1M) pairs in the split of CC12M 1,290… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/pixelprose.imageimage-to-text10M<n<100M173 likes1.2k downloads10mo agoHugging Face05Scaryplasmon96 /PixelArt_Multiview Multiview PixelArt Dataset Summary Contains sets of images representing a full 360° turnaround of characters, animals and objects in pixel art. Each row contains 9 images from all angles. Camera Data can be downloaded Examples Input (f1) f2 f3 f4 f5 f6 f7 f8 f9 Input (f1) f2 f3 f4 f5 f6 f7 f8 f9 Input (f1) f2 f3 f4 f5 f6 f7 f8 f9 Input (f1) f2 f3 f4 f5 f6 f7 f8 f9… See the full description on the dataset page: https://huggingface.co/datasets/Scaryplasmon96/PixelArt_Multiview.imageimage-to-image1K<n<10K4 likes869 downloads1y agoHugging Face06Chan-Y /pixelart-308k PixelArt-308K A large-scale synthetic pixel art image-text dataset containing 308,765 256×256 RGB images, each paired with a text caption describing the image. The dataset was created as a two-stage generation pipeline: Caption generation — prompts were generated using Google's Gemma models through LM Studio. Image generation — the generated prompts were used with FLUX.2-klein-4B to create the corresponding pixel art images. The goal of this dataset is to provide a large… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/pixelart-308k.imagetext-to-image100K<n<1M4 likes803 downloads1mo agoHugging Face07TIGER-Lab /PixelWorld PixelWorld 📜 Paper | 💾 GitHub | 📂 HuggingFace Dataset PixelWorld is a multimodal benchmark that unifies text, tables, code, diagrams, and images into pixel-based inputs (PEAP: Perceive Everything as Pixels). It enables direct comparison between token-based and pixel-based processing. 🔹 Features 📚 Broad Coverage: Text-only (GLUE, SuperGLUE, MMLU-Pro), structured (TableBench), and multimodal tasks (SlidesVQA, WikiSS-QA, MathVerse). 🖼️ Unified Input: Converts… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/PixelWorld.imageany-to-any100K<n<1M6 likes712 downloads2y agoHugging Face08hejun0180-pixel /Wait-Phenomenon-Evidence-Gemini-DeepSeekOriginal Repository: https://huggingface.co/datasets/hejun0180-pixel/Wait-Phenomenon-Evidence-Gemini-DeepSeek Ordinary Agent — Meta-Learning — Autonomous Agent Meta-Learning = Meta-Training = Meta-Social Agent + Civilization Meta-Rules + Emergent Tools WP-AHA: Emergence Tool in LLMs Attributes Cross-Platform:Gemini 1.5 Pro, DeepSeek-V3/R1, Grok, GPT-4o, Claude, Doubao, Qwen, Kimi, Yuanbao. Reproducible: Full Dataset ( 100+ WP-AHA—Endogenous Transition—Raw… See the full description on the dataset page: https://huggingface.co/datasets/hejun0180-pixel/Wait-Phenomenon-Evidence-Gemini-DeepSeek.imagen<1K2 likes567 downloads7mo agoHugging Face09Team-PIXEL /rendered-bookcorpus-bigramsimage1M<n<10M0 likes532 downloads3y agoHugging Face10Obscure-Entropy /PIXELPROSE_HU From Pixels to Prose: A Large Dataset of Dense Image Captions This dataset is an extension of an existing image captioning dataset, enhanced for PixelProse and augmented with Hungarian translations. It provides a valuable resource for researchers and developers working on image captioning, especially those interested in PixelProse and cross-lingual applications. 🌐 Dataset Statistics We report below the number of successfully fetched images and the number of… See the full description on the dataset page: https://huggingface.co/datasets/Obscure-Entropy/PIXELPROSE_HU.imageimage-to-text10M<n<100M5 likes516 downloads2y agoHugging Face11PixelAI-Team /TalkBody4Dgated TalkBody4D Dataset This dataset contains four multi-view image sequences used in our paper "TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting". They are captured with 59 well-calibrated RGB cameras in 20 fps, with a resolution of 3000×4000 and lengths ranging from 800 to 1000 frames. We use the data to evaluate our method for building animatable human body avatars. We also provide the SMPL-X fitting in the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/PixelAI-Team/TalkBody4D.image1M<n<10M100 likes476 downloads2y agoHugging Face12unstonio /pixelgpt-24x24-20k PixelGPT 24×24 — 20K 20,000 native 24×24 pixel-art sprites with captions and semantic taxonomy labels. This is a clean, rebalanced, rights-conscious public subset of the larger PixelGPT 24×24 dataset. Every sprite: is rendered at a native resolution of 24×24 pixels uses no more than 5 colors includes an original text caption is assigned to a two-level semantic taxonomy is distributed in lossless PNG and Parquet formats Looking for the complete dataset?The full edition… See the full description on the dataset page: https://huggingface.co/datasets/unstonio/pixelgpt-24x24-20k.imagetext-to-image10K<n<100K69 likes458 downloads2mo agoHugging Face13Nadav /pixel_squad Dataset Card for "pixel_squad" More Information needed image1M<n<10M0 likes438 downloads3y agoHugging Face14Pixel-Linguist /rendered-sts17 Dataset Summary This dataset is rendered to images from STS-17. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images. Examples of Use Load Arabic to Arabic dataset: from datasets import load_dataset dataset = load_dataset("Pixel-Linguist/rendered-sts17", name="ar-ar", split="test") Load French to English dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts17.image10K<n<100K0 likes407 downloads2y agoHugging Face15Pixel-Linguist /rendered-stsb Dataset Summary This dataset is rendered to images from STS-benchmark. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images. Examples of Use Load English train Dataset: from datasets import load_dataset dataset = load_dataset("Pixel-Linguist/rendered-stsb", name="en", split="train") Load Chinese dev Dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-stsb.imagetext-classification100K<n<1M3 likes396 downloads2y agoHugging Face16Tsomaros /ImageNet-C-pixelate-severity_5image10K<n<100K0 likes385 downloads2y agoHugging Face17physicalai-bmi /forge-arm-pixels physicalai-bmi/forge-arm-pixels Real MuJoCo pixels captured live from the Institute's in-browser Forge arm (WebGPU), paired with the action the released state-checkpoint took. This is the exact training set behind physicalai-bmi/nano-vla-pixels. 2,500 frames across 128 reaches, frames/f#####.png (the rendered MuJoCo arm, 844×520). meta.json — per-frame { i, act:[3], obs:[7], reaches }; act is the 3-D joint-delta action, reaches is the episode index (use it for an episode-level… See the full description on the dataset page: https://huggingface.co/datasets/physicalai-bmi/forge-arm-pixels.imagerobotics1K<n<10K0 likes363 downloads3mo agoHugging Face18DivisonOfficer /Pixel-aligned_RGB-NIR_stereo_dataset Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision CVPR 2025Jinnyeong Kim, Seung-Hwan BaekPOSTECH[arXiv] • [Code] • [Video] • [Dataset on HuggingFace] Overview This repository provides the code and dataset accompanying our CVPR 2025 paper: "Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision" We propose a novel robotic vision system equipped with two pixel-aligned RGB-NIR stereo cameras and a LiDAR sensor mounted on a mobile robot. Our… See the full description on the dataset page: https://huggingface.co/datasets/DivisonOfficer/Pixel-aligned_RGB-NIR_stereo_dataset.image10K<n<100K1 likes321 downloads2y agoHugging Face19PaintBench /pixelsimage1K<n<10K0 likes315 downloads7mo agoHugging Face20TIGER-Lab /PixelReasoner-SFT-DataOverview. The SFT data for training Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning, The queries require fine-grained visual analysis in both images (e.g., infographics, visually-rich scenes, etc) and videos. Details. The data contains 8,000+ reasoning trajectories, including : 2,000+ textual reasoning trajectories, rejection sampled from the base model Qwen2.5-VL-Instruct. These data aims to preserve textual reasoning ability on easier VL… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/PixelReasoner-SFT-Data.imageimage-text-to-text1K<n<10K5 likes294 downloads1y agoHugging Face21Omarrran /Persian_Pixelgated Persian Pixel Persian Pixel is a synthetic optical character recognition (OCR) dataset for Persian / Farsi (fa), in which Unicode text is rendered to images and paired with its exact transcription. It is built for OCR recognition, image-to-text modeling, fine-tuning, and evaluation workflows that need clean, controllable image/label pairs at scale. Because the text is rendered programmatically, every image ships with a perfectly aligned ground-truth label — making the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Persian_Pixel.imageimage-to-text100K<n<1M3 likes283 downloads4mo agoHugging Face22lodestones /pixelprose From Pixels to Prose: A Large Dataset of Dense Image Captions [[ arXiv paper ]] PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions, leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions. @article{pixelprose24, title = {{From Pixels to Prose: A Large Dataset of Dense Image Captions}}, author = {Vasu Singla and Kaiyu Yue and Sukriti Paul and Reza Shirkavand and Mayuka Jayawardhana… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/pixelprose.imageimage-to-text10M<n<100M4 likes260 downloads2y agoHugging Face23isp-uv-es /magicbathynet-s2-pixel-class-taco MagicBathyNet s2 (pixel class) This is a repackaging, not a new dataset. It is MagicBathyNet s2 (pixel class) by TU Berlin RSiM / BIFOLD (Agrafiotis et al.), Dep. of Land and Surveys of Cyprus (LiDAR reference), converted to TACO. Pixel values and labels are kept as released except where the description below says otherwise. All credit belongs to the original authors: if you use it, please cite them and follow their licence. original dataset · paper · licence: CC-BY-NC-4.0… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/magicbathynet-s2-pixel-class-taco.imageimage-segmentation1K<n<10K0 likes228 downloads2d agoHugging Face24Team-PIXEL /rendered-wiki_en-bigramsimage10M<n<100M0 likes207 downloads3y agoHugging Face25quincyu /trace_pixel_v1image100K<n<1M0 likes197 downloads9mo agoHugging Face26carlosuperb /lpc-4view-pixel-art-diffusion LPC 4-View Pixel Art Diffusion Dataset This dataset provides LPC-style pixel-art character sprites for training diffusion models for unconditional and text-to-image character generation. It is developed as part of an undergraduate Third Year Project and is intended for research and educational use. The accompanying training code, preprocessing scripts, and experiments are available in the project GitHub repository: PIXEL-T2I. Dataset Overview The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/carlosuperb/lpc-4view-pixel-art-diffusion.imageunconditional-image-generation10K<n<100K2 likes186 downloads10mo agoHugging Face27to-be /animated-gifs-pixel-art-captions-1k Animated Gifs Pixel Art Captions (1K Sample) A curated sample of 1000 small, square, animated pixel-art gifs, each paired with a natural-language description, broad category labels, and freeform tags -- all generated through an LLM-based enrichment pass by the dataset curator. The gifs themselves were collected online from public osources; the unique value of this dataset is the enrichment layer on top, not the raw images. This is a sample of a much larger corpus (~200,000 gifs)… See the full description on the dataset page: https://huggingface.co/datasets/to-be/animated-gifs-pixel-art-captions-1k.imageimage-to-text1K<n<10K0 likes184 downloads16d agoHugging Face28jiovine /pixel-art-nouns Dataset Card for "pixel-art-nouns" More Information needed image10K<n<100K5 likes183 downloads3y agoHugging Face29jainr3 /diffusiondb-pixelartDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users. The unprecedented scale and diversity of this human-actuated dataset provide exciting research opportunities in understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools to help users more easily use these models.imagetext-to-image1M<n<10M61 likes169 downloads3y agoHugging Face30gigant /pixelprose-bmp-jpg-48image10M<n<100M0 likes167 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.