datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.tiktok-videos-4b
Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies.
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tiktok-videos-4b-dataset
Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies.
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.… See the full description on the dataset page: https://huggingface.co/datasets/0xSojalSec/tiktok-videos-4b-dataset.MT-Video-Bench
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
✨ Introduction
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, existing evaluation benchmarks remain limited to single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios.
🎬 MT-Video-Bench fills this… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/MT-Video-Bench.video-scissors-sessions
Coding agent session traces for kaofelix/video-scissors-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tiktok-videos-sea
TikTok Videos Southeast Asia (MY / TH / ID / SG)
Country-filtered subset of kuben-developer/tiktok-videos-4b:
every row whose country field is MY (Malaysia), TH (Thailand), ID (Indonesia) or SG (Singapore).
Extracted 2026-09-07 from all 27 source parquet files; schema unchanged.
Sizes (verified against the uploaded files)
Config
Country
Rows
Distinct content_id
Ads (is_ad=1)
With caption
my
Malaysia
11,729,283
11,729,283
967,782
9,787,874
th
Thailand… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/tiktok-videos-sea.omniVDCBench
🎬 OmniVDCBench
A bilingual video-domain comprehension benchmark for evaluating fine-grained event understanding, cultural-scene reasoning, and multiple-choice video QA
Dataset Card · Video Question Answering · Fine-grained Event Reasoning · Intangible Cultural Heritage · Chinese / English
Figure 1. Illustration of the OmniVDCBench evaluation setting. The example combines video evidence, multiple-choice questions, model predictions, and category-level… See the full description on the dataset page: https://huggingface.co/datasets/videobenchlab/omniVDCBench.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/grimboy/tiktok-videos-4b.class-to-video-prompt-generationvideo-recs-describe-what-you-seeMore information coming soon
FFmpeg processing script
# Download the list of URLs recursively using wget
cat video_ids_list.txt | parallel -j10 --line-buffer '
# Extract the video name and prepare an output directory
video_name="https://recsys.westlake.edu.cn/MicroLens-100k-Dataset/MicroLens-100k_videos/{}"
output_file="sampled_videos/{}"
# Skip processing if the output file already exists
if [ ! -f "$output_file" ]; then
axel -n 10 --quiet "$video_name"… See the full description on the dataset page: https://huggingface.co/datasets/marcodena/video-recs-describe-what-you-see.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/utopiar/tiktok-videos-4b.VideoScienceBench
VideoScienceBench
A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Dataset Summary
Attribute
Value
Examples
160
Domains
Physics, Chemistry
Format
JSONL (prompt + expected phenomenon + vid)
Data Creation Pipeline
Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kunalmiind/tiktok-videos-4b.qwen3-omni-pairwise-video-infer
Qwen3-Omni Pairwise Video Inference / Evaluation
Pairwise audio-video preference evaluation data for Qwen3-Omni models.
Each sample compares two generated videos (with audio) against a text caption and human/Gemini labels.
Source path on cluster: /inspire/hdd/project/autoregressive-video-generation/public/hym/data/final_infer
Upload snapshot: 2026-06-12 10:45 UTC
Repository layout
Contents of final_infer are uploaded to the dataset repo root:
.cache/
easy500/ —… See the full description on the dataset page: https://huggingface.co/datasets/YinmingHuang/qwen3-omni-pairwise-video-infer.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/0xd4t4/tiktok-videos-4b.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like.
Intended purpose
This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach.
What it shows
clip-level metadata
short and long captions
OCR text
transcript
speaker attribution
structured JSON targets
distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.
