Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RongWei-at /3dfront_render_viewsimage1K<n<10K0 likes36k downloads5mo agoHugging Face02wendlerc /RenderedTextThis dataset has been created by Stability AI and LAION. This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions. Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.imagetext-to-image10M<n<100M59 likes19k downloads1y agoHugging Face03RongWei-at /3dfront-render-viewsimage1K<n<10K0 likes7.5k downloads5mo agoHugging Face04shhdwi /olmocr-pre-rendered olmOCR-bench Pre-Rendered Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model. What This Is The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png(). This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.image1K<n<10K0 likes5.3k downloads7mo agoHugging Face05nateraw /rendered-sst2 Rendered SST-2 The Rendered SST-2 Dataset from Open AI. Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset. This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.imageimage-classification1K<n<10K0 likes3.5k downloads4y agoHugging Face06RongWei-at /3dfront-render-diffuseimage1K<n<10K0 likes3.1k downloads5mo agoHugging Face07RongWei-at /3dfront_renderimage0 likes2.3k downloads5mo agoHugging Face08rodriguescarson /eligible-scroll-atlas-renders Get one mesh in about twenty seconds curl -sO https://raw.githubusercontent.com/rodriguescarson/eligible-scroll-atlas/main/scripts/atlas.py python atlas.py list --ink-pass # the 5 meshes that pass the pre-registered screen python atlas.py ink PHerc0125 z10544_w020 --preview # a downsampled ink map, about 12 KB python atlas.py get PHerc0125 z10544_w020 # the surface volume, 31 planes, plane 15 is the surface from atlas import meshes… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/eligible-scroll-atlas-renders.imagen<1K0 likes2.1k downloads19d agoHugging Face09krahets /dna_rendering_processedgated DNA-Rendering-Processed Dataset Project Page | Paper | Code | Model To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons. To promote future research in the field of human-centric 3D/4D generation, we have open-sourced our re-annotated labels for the DNA-Rendering dataset in this repo, which includes… See the full description on the dataset page: https://huggingface.co/datasets/krahets/dna_rendering_processed.imageimage-to-3d1K<n<10K9 likes1.3k downloads11mo agoHugging Face10eternity304 /procgen-renderformer Procgen RenderFormer Dataset Procedurally generated indoor scenes with ground-truth path-traced renders and precomputed 3-slat VAE latents, built for training RenderFormer-style neural renderers. Each sample is one scene observed from 14 camera poses along an orbit. Configs Config Scenes Samples (scene x frame) Notes main ~307,000 ~4.3 M primary training set zoom ~84,000 ~1.2 M tighter framing variant validation ~1,000 ~14 K held-out assets, not… See the full description on the dataset page: https://huggingface.co/datasets/eternity304/procgen-renderformer.3dimage-to-image1M<n<10M0 likes667 downloads4d agoHugging Face11gt-free-ocr-metrics /omnidocbench-render-compare OmniDocBench Render-and-Compare This dataset contains the rendered HTML reconstructions and comparison images produced by a render-and-compare pipeline — a reference-free visual similarity evaluation framework for OCR systems. Overview The pipeline processes each page of OmniDocBench through a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML (reconstructed.png), and compares it against the original page scan (masked_original.png) using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.imageother10K<n<100K0 likes625 downloads5mo agoHugging Face12Team-PIXEL /rendered-bookcorpus-bigramsimage1M<n<10M0 likes532 downloads3y agoHugging Face13ShapeSplats /Objaverse_2d_rendersgated 2D Image/Depth Rendering of Objaverse Dataset In total, the rendered split contains 167,857 objects. The object ids are in the completed_renders.txt file. After unzipping, the image/depth renders are in the following folder strunture: # e.g., 000-000/000074a334c541878360457c672b6c2e ├── depth.zip ├── image.zip ├── metadata.json └── transforms_train.json Camera Intrinsics 72 views per-object, uniformly sampled on the upper hemisphere Image dimensions: 400×400… See the full description on the dataset page: https://huggingface.co/datasets/ShapeSplats/Objaverse_2d_renders.image100K<n<1M0 likes482 downloads1y agoHugging Face14EntVista /metrixel-character-renders Metrixel Animated Character Renders Multi-view renders, signed-distance-field volumes and per-view mesh tensors produced by Metrixel from a small set of rigged, animated humanoid characters. Each character is captured from four camera angles (0°, 90°, 180°, 270°) across ~30 sampled frames of its motion clip, at 512×512. Every frame/angle carries a matching 64³ signed-distance-field volume and a per-view mesh tensor, so the geometry, the implicit surface and the image are aligned… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-character-renders.imagen<1K0 likes421 downloads22d agoHugging Face15Pixel-Linguist /rendered-sts17 Dataset Summary This dataset is rendered to images from STS-17. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images. Examples of Use Load Arabic to Arabic dataset: from datasets import load_dataset dataset = load_dataset("Pixel-Linguist/rendered-sts17", name="ar-ar", split="test") Load French to English dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts17.image10K<n<100K0 likes407 downloads2y agoHugging Face16Pixel-Linguist /rendered-stsb Dataset Summary This dataset is rendered to images from STS-benchmark. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images. Examples of Use Load English train Dataset: from datasets import load_dataset dataset = load_dataset("Pixel-Linguist/rendered-stsb", name="en", split="train") Load Chinese dev Dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-stsb.imagetext-classification100K<n<1M3 likes396 downloads2y agoHugging Face17Groosezzz /rendered-bookcorpus-8x8image10M<n<100M0 likes395 downloads3y agoHugging Face18Achtik /renderimage1K<n<10K0 likes372 downloads1y agoHugging Face19Caesarrr /objaverse-1.0-renderingsimage1M<n<10M0 likes363 downloads6mo agoHugging Face20RyanWW /audiobench_rendertextaudio10K<n<100K0 likes351 downloads1y agoHugging Face21yianW /spray-drill-renders Spray Bottle + Drill Grasps Functional grasp renders + perception outputs for spray bottles and drills. Previously stored under yianW/mug-grasp-renders/spray_drill_renders/; split out for clarity. Layout <object>_rot<zzz>/ — per-object-rotation case folders Stage 1 renders: image.png, depth.npy, seg.npy, mask_object.npy, cam_pose.npy, intrinsic_K.npy, ... grasp_prompts.json (cases with grasps) grasp_00/..grasp_04/ — per-grasp perception outputs (image_grasp.png, hand… See the full description on the dataset page: https://huggingface.co/datasets/yianW/spray-drill-renders.3dn<1K0 likes282 downloads5mo agoHugging Face22chibifire /anny-render-corpus-generated-train anny-render-corpus-generated Images generated by OmniGen2 from the constructed renders in chibifire/anny-render-corpus. Code: weftspun/anny-render-corpus, on the 6-datasource side of the hexagon. Why this is a separate repository These are generated synthetic, not constructed. They were sampled from a model rather than rendered deterministically from a rig, so their labels are inferred and not true by construction. Our working agreement requires generated data to… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/anny-render-corpus-generated-train.imagen<1K0 likes275 downloads1mo agoHugging Face23DaiPatrick /deadline-render-simulation-20260605152225 Deadline render simulation 20260605152225 Generated by simulate-deadline-render-result.ts This dataset mirrors public data-pack render outputs from Physicl. imagen<1K0 likes243 downloads4mo agoHugging Face24RodelaG /gsm8k-rendered-vlm Rendered GSM8K-VL Dataset Rendered GSM8K-VL is a multimodal math-reasoning dataset for vision-language model evaluation.Each example links: a GSM8K word problem (question) the final numeric answer (answer) cleaned chain-of-thought style reasoning (reasoning) a rendered image path (image) This dataset is intended for controlled experiments comparing text-only and image-based reasoning behavior. Canonical Dataset Artifact The official dataset release uses:… See the full description on the dataset page: https://huggingface.co/datasets/RodelaG/gsm8k-rendered-vlm.imagequestion-answering1K<n<10K0 likes242 downloads5mo agoHugging Face25cminst /transcoda-rendered-row-343k-full-pipeline-v1 Transcoda Rendered Row 343k Full Pipeline v1 Full-page rendered Transcoda row dataset generated from synthetic and random-notation transcriptions. Target contents: 343113 Accepted contents: 343027 Failed/dropped contents: 86 Renderings per accepted content: 4 Accepted images: 1372108 Source counts: {'random': 99990, 'synth': 243037} Staging repo: cminst/transcoda-rendered-row-343k-full-pipeline-v1-shards Each row contains one transcription and four independently rendered page… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-rendered-row-343k-full-pipeline-v1.image100K<n<1M0 likes229 downloads2mo agoHugging Face26Team-PIXEL /rendered-wiki_en-bigramsimage10M<n<100M0 likes207 downloads3y agoHugging Face27clip-benchmark /wds_renderedsst2image1K<n<10K0 likes189 downloads4y agoHugging Face28mteb /rendered-sts13 Dataset Summary This dataset is rendered to images from STS-13. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images. Examples of Use Load test split: from datasets import load_dataset dataset = load_dataset("Pixel-Linguist/rendered-sts13", split="test") Languages English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts13.image1K<n<10K0 likes184 downloads8mo agoHugging Face29tyhuang /ShapeNet_Renderingimage100K<n<1M1 likes174 downloads4y agoHugging Face30ciderlab /amex_render_pairs_v3image10K<n<100K0 likes165 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.