Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tals /vitaminc Details Fact Verification dataset created for Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence (Schuster et al., NAACL 21`) based on Wikipedia edits (revisions). For more details see: https://github.com/TalSchuster/VitaminC When using this dataset, please cite the paper: BibTeX entry and citation info @inproceedings{schuster-etal-2021-get, title = "Get Your Vitamin {C}! Robust Fact Verification with Contrastive Evidence", author =… See the full description on the dataset page: https://huggingface.co/datasets/tals/vitaminc.texttext-classification100K<n<1M11 likes4.7k downloads4y agoHugging Face02Emanresu /features-dinov3-vith16plus-224-imagenet-22k-wdstext1M<n<10M0 likes2.1k downloads11mo agoHugging Face03MIT-Media-Lab /oakink2-vitra-streaming-v1 OakInk2 → VITRA Stage-1 (complete audited release) This repository contains all 627 physical OakInk2-TaMF sequences converted to VITRA Stage-1. Every sequence source pair is pinned to kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact frame identity, converted across the four calibrated views, checked by geometry and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded, and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.imagevideo-classificationn<1K0 likes873 downloads2mo agoHugging Face04ViTeX-Bench /ViTeX-Dataset ViTeX-Dataset 📄 Paper &nbsp;·&nbsp; 🌐 Project page &nbsp;·&nbsp; 📊 Dataset &nbsp;·&nbsp; 🧪 Code &nbsp;·&nbsp; 🤖 Model weights &nbsp;·&nbsp; 🏆 Leaderboard Paired real-video dataset for video scene text editing: given a source video, a binary text-region mask, and a (source string → target string) pair, replace only the masked scene text across all frames while preserving the rest of the scene. Accepted to NeurIPS 2026 E&D Track. Authors: Xinghao Chen, Xiangbo Gao, Jiongze… See the full description on the dataset page: https://huggingface.co/datasets/ViTeX-Bench/ViTeX-Dataset.textvideo-to-videon<1K0 likes729 downloads5d agoHugging Face05vitaliy-sharandin /energy-consumption-hourly-spaintabular10K<n<100K3 likes607 downloads3y agoHugging Face06VItaldob /viciebski-pralietaryj-yiddish Viciebski Pralietaryj — Yiddish blocks Blocks of newspaper text set in Yiddish (Hebrew script), cut from scans of Viciebski pralietaryj («Віцебскі пралетарый»), a newspaper published in Vitebsk, Byelorussian SSR (Belarus), in 1930, 1931 and 1933. 2684 images from 78 pages across 78 issues. No transcriptions — this is a raw corpus for OCR/HTR work, not a labelled set. Belarusian blocks are present. The Yiddish pages ran as inserts inside a Belarusian newspaper, and blocks were… See the full description on the dataset page: https://huggingface.co/datasets/VItaldob/viciebski-pralietaryj-yiddish.imageimage-to-text1K<n<10K0 likes496 downloads1mo agoHugging Face07Qdrant /wolt-food-clip-ViT-B-32-embeddings wolt-food-clip-ViT-B-32-embeddings Qdrant's Food Discovery demo relies on the dataset of food images from the Wolt app. Each point in the collection represents a dish with a single image. The image is represented as a vector of 512 float numbers. Generation process The embeddings generated with clip-ViT-B-32 model have been generated using the following code snippet: from PIL import Image from sentence_transformers import SentenceTransformer image_path =… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings.imagefeature-extraction1M<n<10M9 likes487 downloads3y agoHugging Face08benikm91 /sketch-graph-vitruvion pretty_name: SketchGraphs, Vitruvion selection (rebuilt) license: other license_name: onshape-terms-of-use license_link: https://www.onshape.com/legal/terms-of-use#your_content size_categories: - 1M<n<10M tags: - cad - parametric-cad - sketches - geometric-constraints - sketchgraphs - vitruvion SketchGraphs, Vitruvion selection (sg_filtered_unique.npy, rebuilt) This is the dataset of Vitruvion (Seff et al., Vitruvion: A Generative Model of… See the full description on the dataset page: https://huggingface.co/datasets/benikm91/sketch-graph-vitruvion.text1M<n<10M0 likes444 downloads5d agoHugging Face09epfl-vita /svi-benchmark Stable Video Infinity (SVI) Benchmark Dataset This benchmark dataset is introduced in the paper: Stable Video Infinity: Infinite-Length Video Generation with Error Recycling by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025). Project page: https://stable-video-infinity.github.io/homepage/ Code: https://github.com/vita-epfl/Stable-Video-Infinity Abstract We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.imageimage-to-videon<1K7 likes430 downloads1y agoHugging Face10closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes415 downloads4y agoHugging Face11duckking032 /vital Vital PresetShare Renders Rendered Vital presets scraped from PresetShare. sample rate: 22050 render duration: 6.0s MIDI note: 72 (C4) note duration: 5.0s velocity: 100 Files are organized under by_type/<sound-type>/<preset-id>_<name>/ with: preset.vital preview.mp3 vital-render.wav metadata.json See manifest.jsonl and summary.json for run metadata. audion<1K0 likes372 downloads3mo agoHugging Face12grspo /sam2-vit-btabularn<1K0 likes316 downloads11mo agoHugging Face13closji /flickr30k_clip-ViT-B-32-caption_pairstabular10M<n<100M4 likes314 downloads4y agoHugging Face14pixxu /ViTextRender-500K Vietnamese Text Render 500K Dataset A large-scale dataset containing 500K Vietnamese text rendering image-text pairs for training generative models to improve text rendering performance. Dataset Structure image: Rendered text image in PNG format text: Corresponding text content filename: Original filename Usage This dataset is designed for fine-tuning generative models to improve text rendering capabilities on Vietnamese language. from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/pixxu/ViTextRender-500K.texttext-to-image100K<n<1M0 likes312 downloads10mo agoHugging Face15forgeml /viton_hd Dataset Card for "viton_hd" More Information needed image10K<n<100K11 likes305 downloads3y agoHugging Face16mekaneeky /Synthetic_English_VITS_22.5k Dataset Card for "Synthetic_English_VITS_22.5k" More Information needed text1K<n<10K1 likes257 downloads3y agoHugging Face17lscpku /VITATECSgated Dataset Card for VITATECS Dataset Description Dataset Summary VITATECS is a diagnostic VIdeo-Text dAtaset for the evaluation of TEmporal Concept underStanding. [2023/11/27] We have updated a new version of VITATECS which is generated using ChatGPT. The previous version generated by OPT-175B can be found here. Languages English. Dataset Structure Usage aspect = 'Type' #… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/VITATECS.text10K<n<100K5 likes253 downloads2y agoHugging Face18nguyensu27 /VITS_DATASET_60k_100Kaudio100K<n<1M0 likes251 downloads9mo agoHugging Face19Martingkc /LLaVa-CC3M-Pretrain-clip-vit-base-patch32text100K<n<1M0 likes248 downloads6mo agoHugging Face20grspo /sam2-vit-stabularn<1K0 likes240 downloads11mo agoHugging Face21Martingkc /LLaVa-CC3M-PostTraining-clip-vit-base-patch16text100K<n<1M0 likes236 downloads6mo agoHugging Face22leungtianle /new-rl-vitaaudio100K<n<1M2 likes225 downloads10mo agoHugging Face23vitaliy-sharandin /energy-consumption-weather-hourly-spaintabular100K<n<1M4 likes216 downloads3y agoHugging Face24meituan-longcat /VitaBench-2.0textn<1K8 likes204 downloads4mo agoHugging Face25V4ldeLund /vital-articles-da-wiki Vital Articles Danish Wikipedia Dataset Overview Total articles: 28,006 Files: 29 Parquet shards Language: Danish (da) Contents Each row is one article with these fields: en_title: English Wikipedia title da_title: Danish Wikipedia title da_url: Danish Wikipedia article URL markdown: Article content in markdown format markdown_chars: Character count of markdown source_lang: Source language code (da) fetched_at_utc: UTC timestamp when the article was fetched texttext-generation10K<n<100K0 likes199 downloads8mo agoHugging Face26Emanresu /cls_dinov3-vith16plus_in22ktabulartoken-classification100K<n<1M0 likes184 downloads11mo agoHugging Face27jablonkagroup /sarscov2_vitro_touret Dataset Details Dataset Description An in-vitro screen of the Prestwick chemical library composed of 1,480 approved drugs in an infected cell-based assay. Curated by: License: CC BY 4.0 Dataset Sources corresponding publication Data source Citation BibTeX: @article{Touret2020, doi = {10.1038/s41598-020-70143-6}, url = {https://doi.org/10.1038/s41598-020-70143-6}, year = {2020}, month = aug, publisher = {Springer Science and Business Media LLC}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/sarscov2_vitro_touret.tabular10K<n<100K0 likes179 downloads1y agoHugging Face28closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes169 downloads4y agoHugging Face29fzhu22 /imagenet1k-vit-preproc-4text100K<n<1M0 likes169 downloads1y agoHugging Face30VITA-MLLM /Comic-9K Comic-9K Image Extracting all images. cat images.tar.gz.aa images.tar.gz.ab images.tar.gz.ac images.tar.gz.ad images.tar.gz.ae > images.tar.gz tar xvzf images.tar.gz Summary We provide human-written plot synopsis. summary.jsonl image100K<n<1M6 likes156 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.