Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kwakuobeng /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes2.4k downloads1mo agoHugging Face02elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K3 likes2.1k downloads1y agoHugging Face03merway /tiktok-videos-4b Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes1.5k downloads1mo agoHugging Face04blaccastro /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes1.2k downloads1mo agoHugging Face05dams2005 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes746 downloads1mo agoHugging Face06KOM-00 /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes609 downloads1mo agoHugging Face07yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes599 downloads7mo agoHugging Face08alex12223322 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes457 downloads1mo agoHugging Face09ansulev /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes434 downloads1mo agoHugging Face10mrfakename /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes433 downloads1mo agoHugging Face11abdellatifinformation /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes362 downloads1mo agoHugging Face12seanphan /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes361 downloads1mo agoHugging Face13hojj /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes300 downloads1mo agoHugging Face140xSojalSec /tiktok-videos-4b-dataset Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research.… See the full description on the dataset page: https://huggingface.co/datasets/0xSojalSec/tiktok-videos-4b-dataset.tabulartext-classification1B<n<10B0 likes286 downloads17d agoHugging Face15NJU-LINK /MT-Video-Benchgated MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues ✨ Introduction Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, existing evaluation benchmarks remain limited to single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. 🎬 MT-Video-Bench fills this… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/MT-Video-Bench.videotext-generation6 likes228 downloads9mo agoHugging Face16kaofelix /video-scissors-sessions Coding agent session traces for kaofelix/video-scissors-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.tabulartext-generationn<1K0 likes192 downloads6mo agoHugging Face17superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face18kkndlee /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes172 downloads1mo agoHugging Face19beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging Face20khursanirevo /tiktok-videos-sea TikTok Videos Southeast Asia (MY / TH / ID / SG) Country-filtered subset of kuben-developer/tiktok-videos-4b: every row whose country field is MY (Malaysia), TH (Thailand), ID (Indonesia) or SG (Singapore). Extracted 2026-09-07 from all 27 source parquet files; schema unchanged. Sizes (verified against the uploaded files) Config Country Rows Distinct content_id Ads (is_ad=1) With caption my Malaysia 11,729,283 11,729,283 967,782 9,787,874 th Thailand… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/tiktok-videos-sea.tabulartext-classification10M<n<100M0 likes165 downloads1mo agoHugging Face21videobenchlab /omniVDCBench 🎬 OmniVDCBench A bilingual video-domain comprehension benchmark for evaluating fine-grained event understanding, cultural-scene reasoning, and multiple-choice video QA Dataset Card · Video Question Answering · Fine-grained Event Reasoning · Intangible Cultural Heritage · Chinese / English Figure 1. Illustration of the OmniVDCBench evaluation setting. The example combines video evidence, multiple-choice questions, model predictions, and category-level… See the full description on the dataset page: https://huggingface.co/datasets/videobenchlab/omniVDCBench.textvisual-question-answering1K<n<10K1 likes122 downloads3mo agoHugging Face22grimboy /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/grimboy/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes121 downloads1mo agoHugging Face23shreyahegde /class-to-video-prompt-generationtexttext-generationn<1K3 likes88 downloads2y agoHugging Face24marcodena /video-recs-describe-what-you-seeMore information coming soon FFmpeg processing script # Download the list of URLs recursively using wget cat video_ids_list.txt | parallel -j10 --line-buffer ' # Extract the video name and prepare an output directory video_name="https://recsys.westlake.edu.cn/MicroLens-100k-Dataset/MicroLens-100k_videos/{}" output_file="sampled_videos/{}" # Skip processing if the output file already exists if [ ! -f "$output_file" ]; then axel -n 10 --quiet "$video_name"… See the full description on the dataset page: https://huggingface.co/datasets/marcodena/video-recs-describe-what-you-see.texttext-generation10K<n<100K1 likes87 downloads1y agoHugging Face25utopiar /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/utopiar/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes84 downloads25d agoHugging Face26lmgame /VideoScienceBench VideoScienceBench A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon). Dataset Summary Attribute Value Examples 160 Domains Physics, Chemistry Format JSONL (prompt + expected phenomenon + vid) Data Creation Pipeline Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.textquestion-answeringn<1K3 likes56 downloads8mo agoHugging Face27kunalmiind /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/kunalmiind/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes35 downloads1mo agoHugging Face28YinmingHuang /qwen3-omni-pairwise-video-infer Qwen3-Omni Pairwise Video Inference / Evaluation Pairwise audio-video preference evaluation data for Qwen3-Omni models. Each sample compares two generated videos (with audio) against a text caption and human/Gemini labels. Source path on cluster: /inspire/hdd/project/autoregressive-video-generation/public/hym/data/final_infer Upload snapshot: 2026-06-12 10:45 UTC Repository layout Contents of final_infer are uploaded to the dataset repo root: .cache/ easy500/ —… See the full description on the dataset page: https://huggingface.co/datasets/YinmingHuang/qwen3-omni-pairwise-video-infer.videotext-generationn<1K0 likes30 downloads4mo agoHugging Face290xd4t4 /tiktok-videos-4bgated TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/0xd4t4/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes29 downloads1mo agoHugging Face30metavi /video-understanding-distillation-sample Video Understanding Distillation Sample This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like. Intended purpose This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach. What it shows clip-level metadata short and long captions OCR text transcript speaker attribution structured JSON targets distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.texttext-generationn<1K0 likes27 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.