datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.oakink2-vitra-streaming-v1
OakInk2 → VITRA Stage-1 (complete audited release)
This repository contains all 627 physical OakInk2-TaMF sequences converted to
VITRA Stage-1. Every sequence source pair is pinned to
kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact
frame identity, converted across the four calibrated views, checked by geometry
and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded,
and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.StreamingBenchStreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.streamingvlm-task-aware-vqa-suite-rowwise
StreamingVLM Task-Aware VQA Suite (Row-wise)
Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row.
Task groups
task_group
Dataset configs
coarse_object_presence
pope, repope, hpope
general_scene_understanding
vqav2, gqa
fine_grained_visual_evidence
gqa_attribute, mmbench
textual_fine_grained_evidence
textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.StreamingOmniDatasets
StreamingOmniDatasets v0.5.0 Streaming CoT Mixed
Four training-ready configs with slim viewer schemas. Structured CoT stores concise, auditable causal state updates rather than private model thinking. Full duplex, barge-in, and simultaneous listen/speak are intentionally deferred.
conversational-streaming-asr-benchmark
SquadStack Conversational Streaming ASR Benchmark (8 kHz)
Version 1.0.0 · maintained by SquadStack
Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use
Key takeaways
What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.
