Team Ai
Datasetpublic

MCG-NJU/VideoChat3-Training-Data-Annotations

VideoChat3-Stage3-Training-Data This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3. You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry. To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
2likes654downloads
Dataset Card

VideoChat3-Stage3-Training-Data

This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3.

You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry.

To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected and curated, we additionally provide a more detailed introduction at VideoChat3-Stage3-Training-Data.

📄 Paper · 🌐 Homepage · 💻 GitHub · 🤗 Paper Page

<p align="center"> <img src="training_pipeline.jpg" alt="VideoChat3 Training Overview" width="100%"> </p>

Data Organization

ComponentDescription
training_data_annotationsTraining annotation data for each stage
videochat3_data_annotationsAnnotation files for individual datasets
training_data_source_tablesSource datasets and paths

In videochat3_data_annotations, the Stage 0 and Stage 2 annotation data partially overlap with our previously released VideoChat3-Academic2M and VideoChat3-LV116K. For convenience of downloading and use, we re-upload them here.

[!Note] Because the training_data_source_tables contain data from a wide range of sources, there may be occasional errors. If you spot any issues, please feel free to contact me at zzhiqiu997@gmail.com.

Citation

If you use this data, please cite VideoChat3 and the original video datasets used by the annotations.

@article{li2026videochat3,
  title={VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding},
  author={Li, Xinhao and Zhu, Yuhan and Zeng, Xiangyu and Dong, Yuhao and Wu, Haoning and Zhang, Zhiqiu and Yang, Yuandong and Ma, Changlian and Zhang, Qingyu and Shi, Yansong and others},
  journal={arXiv preprint arXiv:2607.14935},
  year={2026}
}