anonymsubs/postvalid-v2-test-set
LongShOTBench (Test Split) Benchmark accompanying the NeurIPS 2026 submission "A Benchmark for Omni-Modal Reasoning in Long Videos." This dataset is shared anonymously to support double-blind review. Purpose LongShOTBench evaluates multimodal LLMs on long-form video understanding across vision, speech, and non-speech audio, using intent-driven questions and weighted criterion-level rubrics. Intended for evaluation only, not training. License CC… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.
LongShOTBench (Test Split)
Benchmark accompanying the NeurIPS 2026 submission "A Benchmark for Omni-Modal Reasoning in Long Videos."
This dataset is shared anonymously to support double-blind review.
Purpose
LongShOTBench evaluates multimodal LLMs on long-form video understanding across vision, speech, and non-speech audio, using intent-driven questions and weighted criterion-level rubrics. Intended for evaluation only, not training.
License
CC BY-NC-SA 4.0. Source videos are referenced by ID and not redistributed.
