OpenGVLab/MVBench
MVBench Important Update [18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference. We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.
4820k
1---2license: mit3extra_gated_prompt: >-4 You agree to not use the dataset to conduct experiments that cause harm to5 human subjects. Please note that the data in this dataset may be subject to6 other agreements. Before using the data, be sure to read the relevant7 agreements carefully to ensure compliant use. Video copyrights belong to the8 original video creators or platforms and are for academic research use only.9task_categories:10- visual-question-answering11- video-classification12extra_gated_fields:13 Name: text14 Company/Organization: text15 Country: text16 E-Mail: text17modalities:18- Video19- Text20configs:21- config_name: action_sequence22 data_files: json/action_sequence.json23- config_name: moving_count24 data_files: json/moving_count.json25- config_name: action_prediction26 data_files: json/action_prediction.json27- config_name: episodic_reasoning28 data_files: json/episodic_reasoning.json29- config_name: action_antonym30 data_files: json/action_antonym.json31- config_name: action_count32 data_files: json/action_count.json33- config_name: scene_transition34 data_files: json/scene_transition.json35- config_name: object_shuffle36 data_files: json/object_shuffle.json37- config_name: object_existence38 data_files: json/object_existence.json39- config_name: fine_grained_pose40 data_files: json/fine_grained_pose.json41- config_name: unexpected_action42 data_files: json/unexpected_action.json43- config_name: moving_direction44 data_files: json/moving_direction.json45- config_name: state_change46 data_files: json/state_change.json47- config_name: object_interaction48 data_files: json/object_interaction.json49- config_name: character_order50 data_files: json/character_order.json51- config_name: action_localization52 data_files: json/action_localization.json53- config_name: counterfactual_inference54 data_files: json/counterfactual_inference.json55- config_name: fine_grained_action56 data_files: json/fine_grained_action.json57- config_name: moving_attribute58 data_files: json/moving_attribute.json59- config_name: egocentric_navigation60 data_files: json/egocentric_navigation.json61language:62- en63size_categories:64- 1K<n<10K65---66# MVBench67 68## Dataset Description69 70- **Repository:** [MVBench](https://github.com/OpenGVLab/Ask-Anything/blob/main/video_chat2/mvbench.ipynb)71- **Paper:** [2311.17005](https://arxiv.org/abs/2311.17005)72- **Point of Contact:** mailto:[kunchang li](likunchang@pjlab.org.cn)73 74 75## <span style="color: red;">Important Update</span>76[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit [ROSE Lab](https://rose1.ntu.edu.sg/dataset/actionRecognition/) to access the data. We also provide a [list of the 320 videos](https://huggingface.co/datasets/OpenGVLab/MVBench/blob/main/video/MVBench_videos_ntu.txt) used in MVBench for your reference.77 78 7980 81We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from perception to cognition. Guided by task definitions, we then **automatically transform public video annotations into multiple-choice QA** for task evaluation. This unique paradigm enables efficient creation of MVBench with minimal manual intervention while ensuring evaluation fairness through ground-truth video annotations and avoiding biased LLM scoring. The **20** temporal task examples are as follows.82 8384 85## Evaluation86 87An evaluation example is provided in [mvbench.ipynb](https://github.com/OpenGVLab/Ask-Anything/blob/main/video_chat2/mvbench.ipynb). Please follow the pipeline to prepare the evaluation code for various MLLMs.88 89- **Preprocess**: We preserve the raw video (high resolution, long duration, etc.) along with corresponding annotations (start, end, subtitles, etc.) for future exploration; hence, the decoding of some raw videos like Perception Test may be slow.90- **Prompt**: We explore effective system prompts to encourage better temporal reasoning in MLLM, as well as efficient answer prompts for option extraction.91 92## Leadrboard93 94While an [Online leaderboard]() is under construction, the current standings are as follows:95 96