Team Ai
Datasetpublic

zhaochenyang20/Video_AMME_ci

Video-AMME CI Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci. Each example keeps the Video-MME video and moves the question, answer choices, and answer-format instruction into a spoken WAV file. Files data/test.jsonl: metadata and source Video-MME references. audios/*.wav: spoken question/options/instruction. videos/*.mp4: present only when built with --copy-videos. Generation TTS model: fishaudio/s2-pro Max samples… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes45downloads
Dataset Card

Video-AMME CI

Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci. Each example keeps the Video-MME video and moves the question, answer choices, and answer-format instruction into a spoken WAV file.

Files

  • —data/test.jsonl: metadata and source Video-MME references.
  • —audios/*.wav: spoken question/options/instruction.
  • —videos/*.mp4: present only when built with --copy-videos.

Generation

  • —TTS model: fishaudio/s2-pro
  • —Max samples requested: 50
  • —Videos copied into this dataset: False

The benchmark request text must contain routing/format instructions only; the semantic question content lives in audio_text and the WAV file.