datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.Zhoulifeng-Streaming-Dataset
Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线)
本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。
本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。
🗂️ 核心处理流程 (Data Pipeline)
数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明:
Step 01: 语音文本标点恢复 (Punctuation Generation)
脚本: 01_punctuation_generate.py
输入/输出: 生成 01_punc_transcripts/ 目录
说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.oakink2-vitra-streaming-v1
OakInk2 → VITRA Stage-1 (complete audited release)
This repository contains all 627 physical OakInk2-TaMF sequences converted to
VITRA Stage-1. Every sequence source pair is pinned to
kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact
frame identity, converted across the four calibrated views, checked by geometry
and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded,
and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.StreamingBenchvietspeech-train-streamingStreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.combined-dataset-streaming-large-teststreaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.tsa-throughput-streaming-test4tsa-throughput-streaming-test2smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.streamingvlm-task-aware-vqa-suite-rowwise
StreamingVLM Task-Aware VQA Suite (Row-wise)
Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row.
Task groups
task_group
Dataset configs
coarse_object_presence
pope, repope, hpope
general_scene_understanding
vqav2, gqa
fine_grained_visual_evidence
gqa_attribute, mmbench
textual_fine_grained_evidence
textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.combined-dataset-streamingcombined-dataset-streaming-largecombined-dataset-streaming-large-9streamingqatsa-throughput-streaming-test3streamingqa_sanitizedStreamingOmniDatasets
StreamingOmniDatasets v0.5.0 Streaming CoT Mixed
Four training-ready configs with slim viewer schemas. Structured CoT stores concise, auditable causal state updates rather than private model thinking. Full duplex, barge-in, and simultaneous listen/speak are intentionally deferred.
streaming-gebd-causal
Audited Kinetics-GEBD Causal Metadata
This metadata-only release converts the publicly released Kinetics-GEBD
annotations into an auditable 24 FPS causal training representation. It does
not redistribute Kinetics or YouTube video bytes.
Splits
Hub split
Official source file
Records
Meaning
train
k400_train_raw_annotation.pkl
18,808
Public GEBD training annotations
validation
k400_val_raw_annotation.pkl
18,815
Public Kinetics-GEBD validation… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streaming-gebd-causal.conversational-streaming-asr-benchmark
SquadStack Conversational Streaming ASR Benchmark (8 kHz)
Version 1.0.0 · maintained by SquadStack
Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use
Key takeaways
What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.combined-dataset-streaming-large-7combined-dataset-streaming-large-4combined-dataset-streaming-large-5model0-boundary-streaming
Model Card: Streaming Terminal Log Boundary Predictor (Phi-4 LoRA)
🤖 Model Details
Base Model: unsloth/Phi-4-unsloth-bnb-4bit (14B Parameters)
Architecture: LoRA Adapters (PEFT)
Task: Binary Classification (Terminal Event Boundary Detection)
Quantization: 4-bit (bitsandbytes)
Language: English / Bash / Terminal XML
🎯 Intended Use
This model serves as "Model 0" for the Winter 2026 iteration of the AutoDocs project. Its primary function is to segment a… See the full description on the dataset page: https://huggingface.co/datasets/librocubic/model0-boundary-streaming.combined-dataset-streaming-small-testVQA-stage2-streamingtsa-throughput-streaming-testtarteel-ai-EA-DI-combined-stream-streamingstreamingqa-r2l
