Team Ai
Datasetpublic

OpenGVLab/MVBench

MVBench Important Update [18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference. We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.

sourceHugging Facemitupdated 2y agoView on Hugging Face
48likes20kdownloads
README.md96 linesDownload Raw Back to root
1---2license: mit3extra_gated_prompt: >-4  You agree to not use the dataset to conduct experiments that cause harm to5  human subjects. Please note that the data in this dataset may be subject to6  other agreements. Before using the data, be sure to read the relevant7  agreements carefully to ensure compliant use. Video copyrights belong to the8  original video creators or platforms and are for academic research use only.9task_categories:10- visual-question-answering11- video-classification12extra_gated_fields:13  Name: text14  Company/Organization: text15  Country: text16  E-Mail: text17modalities:18- Video19- Text20configs:21- config_name: action_sequence22  data_files: json/action_sequence.json23- config_name: moving_count24  data_files: json/moving_count.json25- config_name: action_prediction26  data_files: json/action_prediction.json27- config_name: episodic_reasoning28  data_files: json/episodic_reasoning.json29- config_name: action_antonym30  data_files: json/action_antonym.json31- config_name: action_count32  data_files: json/action_count.json33- config_name: scene_transition34  data_files: json/scene_transition.json35- config_name: object_shuffle36  data_files: json/object_shuffle.json37- config_name: object_existence38  data_files: json/object_existence.json39- config_name: fine_grained_pose40  data_files: json/fine_grained_pose.json41- config_name: unexpected_action42  data_files: json/unexpected_action.json43- config_name: moving_direction44  data_files: json/moving_direction.json45- config_name: state_change46  data_files: json/state_change.json47- config_name: object_interaction48  data_files: json/object_interaction.json49- config_name: character_order50  data_files: json/character_order.json51- config_name: action_localization52  data_files: json/action_localization.json53- config_name: counterfactual_inference54  data_files: json/counterfactual_inference.json55- config_name: fine_grained_action56  data_files: json/fine_grained_action.json57- config_name: moving_attribute58  data_files: json/moving_attribute.json59- config_name: egocentric_navigation60  data_files: json/egocentric_navigation.json61language:62- en63size_categories:64- 1K<n<10K65---66# MVBench67 68## Dataset Description69 70- **Repository:** [MVBench](https://github.com/OpenGVLab/Ask-Anything/blob/main/video_chat2/mvbench.ipynb)71- **Paper:** [2311.17005](https://arxiv.org/abs/2311.17005)72- **Point of Contact:** mailto:[kunchang li](likunchang@pjlab.org.cn)73 74 75## <span style="color: red;">Important Update</span>76[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit [ROSE Lab](https://rose1.ntu.edu.sg/dataset/actionRecognition/) to access the data. We also provide a [list of the 320 videos](https://huggingface.co/datasets/OpenGVLab/MVBench/blob/main/video/MVBench_videos_ntu.txt) used in MVBench for your reference.77 78 79![images](./assert/generation.png)80 81We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from perception to cognition. Guided by task definitions, we then **automatically transform public video annotations into multiple-choice QA** for task evaluation. This unique paradigm enables efficient creation of MVBench with minimal manual intervention while ensuring evaluation fairness through ground-truth video annotations and avoiding biased LLM scoring. The **20** temporal task examples are as follows.82 83![images](./assert/task_example.png)84 85## Evaluation86 87An evaluation example is provided in [mvbench.ipynb](https://github.com/OpenGVLab/Ask-Anything/blob/main/video_chat2/mvbench.ipynb). Please follow the pipeline to prepare the evaluation code for various MLLMs.88 89- **Preprocess**: We preserve the raw video (high resolution, long duration, etc.) along with corresponding annotations (start, end, subtitles, etc.) for future exploration; hence, the decoding of some raw videos like Perception Test may be slow.90- **Prompt**: We explore effective system prompts to encourage better temporal reasoning in MLLM, as well as efficient answer prompts for option extraction.91 92## Leadrboard93 94While an [Online leaderboard]() is under construction, the current standings are as follows:95 96![images](./assert/leaderboard.png)