Team Ai
Datasetpublic

anonymsubs/postvalid-v2-test-set

LongShOTBench (Test Split) Benchmark accompanying the NeurIPS 2026 submission "A Benchmark for Omni-Modal Reasoning in Long Videos." This dataset is shared anonymously to support double-blind review. Purpose LongShOTBench evaluates multimodal LLMs on long-form video understanding across vision, speech, and non-speech audio, using intent-driven questions and weighted criterion-level rubrics. Intended for evaluation only, not training. License CC… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.

sourceHugging Facecc-by-nc-sa-4.0updated 5mo agoView on Hugging Face
0likes8downloads
Dataset Card

LongShOTBench (Test Split)

Benchmark accompanying the NeurIPS 2026 submission "A Benchmark for Omni-Modal Reasoning in Long Videos."

This dataset is shared anonymously to support double-blind review.

Purpose

LongShOTBench evaluates multimodal LLMs on long-form video understanding across vision, speech, and non-speech audio, using intent-driven questions and weighted criterion-level rubrics. Intended for evaluation only, not training.

License

CC BY-NC-SA 4.0. Source videos are referenced by ID and not redistributed.