Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes11k downloads1y agoHugging Face02hzb29 /Zhoulifeng-Streaming-Dataset Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线) 本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。 本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。 🗂️ 核心处理流程 (Data Pipeline) 数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明: Step 01: 语音文本标点恢复 (Punctuation Generation) 脚本: 01_punctuation_generate.py 输入/输出: 生成 01_punc_transcripts/ 目录 说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.textfeature-extraction1K<n<10K7 likes2.2k downloads7mo agoHugging Face03xavierdurawa /proof-pile-2-streaming ArXiv | Models | Data | Code | Blog | Sample Explorer Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, Sean Welleck The Proof-Pile-2 is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the Llemma 7B and Llemma 34B models. It consists of three subsets: arxiv (29B tokens): the ArXiv subset of RedPajama open-web-math (15B tokens): The OpenWebMath… See the full description on the dataset page: https://huggingface.co/datasets/xavierdurawa/proof-pile-2-streaming.text-generation10B<n<100B1 likes1.7k downloads3y agoHugging Face04xrorrim /streaming_vlmvideo0 likes1.4k downloads1y agoHugging Face05MIT-Media-Lab /oakink2-vitra-streaming-v1 OakInk2 → VITRA Stage-1 (complete audited release) This repository contains all 627 physical OakInk2-TaMF sequences converted to VITRA Stage-1. Every sequence source pair is pinned to kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact frame identity, converted across the four calibrated views, checked by geometry and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded, and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.imagevideo-classificationn<1K0 likes841 downloads2mo agoHugging Face06freesky /streamingbenchvideon<1K0 likes522 downloads11mo agoHugging Face07aiai-laboratory /vietspeech-train-streamingtext100K<n<1M0 likes425 downloads2mo agoHugging Face08ngqtrung /StreamingBenchimage1K<n<10K0 likes352 downloads8mo agoHugging Face09jayzhu486 /StreamingBench-Slice StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.imagequestion-answering1K<n<10K1 likes318 downloads6mo agoHugging Face10RuiLu /combined-dataset-streaming-large-testtext1M<n<10M0 likes286 downloads11mo agoHugging Face11THU-SI /Spatial-TTT-Data-Streaming1 likes266 downloads7mo agoHugging Face12JohnVitz /tsa-throughput-streaming-test4tabular100K<n<1M0 likes253 downloads8mo agoHugging Face13DeliberatorArchiver /hls_streaming_media0 likes243 downloads3y agoHugging Face14JohnVitz /tsa-throughput-streaming-test2tabular10K<n<100K0 likes242 downloads8mo agoHugging Face15sumathiselvan /LAION-Mobile-streaming0 likes236 downloads2mo agoHugging Face16RuiLu /combined-dataset-streamingtext10M<n<100M0 likes218 downloads11mo agoHugging Face17david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes196 downloads4mo agoHugging Face18vkatg /streaming-phi-deidentification-benchmark Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.documentn<1K0 likes187 downloads7mo agoHugging Face19kfkas /streamingvlm-task-aware-vqa-suite-rowwise StreamingVLM Task-Aware VQA Suite (Row-wise) Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row. Task groups task_group Dataset configs coarse_object_presence pope, repope, hpope general_scene_understanding vqav2, gqa fine_grained_visual_evidence gqa_attribute, mmbench textual_fine_grained_evidence textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.image10K<n<100K0 likes180 downloads2mo agoHugging Face20Pramish /egooops_streamingvideon<1K0 likes173 downloads1y agoHugging Face21emgena /omnimcp_data_kafka_streaming_teaser 🚀 DataOps Kafka Streaming & Schema Registry Guard (Evaluation Teaser + Turnkey MCP Server) ⚡ Official Free Community Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Production Master Package on Gumroad:👉 Purchase Full Enterprise Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout! (Starting at €49) ⚡ Activate in Cursor IDE & Claude Desktop in 30 Seconds This repository now contains a zero-dependency, turnkey Model Context Protocol… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_data_kafka_streaming_teaser.n<1K0 likes173 downloads12d agoHugging Face22RuiLu /combined-dataset-streaming-large-9text10M<n<100M0 likes160 downloads10mo agoHugging Face23RuiLu /combined-dataset-streaming-largetext1M<n<10M0 likes139 downloads11mo agoHugging Face24StreamingVLM /streaming_vlm0 likes137 downloads1y agoHugging Face25Efficient-Large-Model /SANA-Streaming-example-training-dataset SANA-Streaming Example Training Dataset This repository contains 1,000 aligned reverse video-editing pairs for the public SANA-Streaming bidirectional V2V training recipe. License and Terms This dataset is made available for non-commercial research use only under the terms in LICENSE. See NOTICE.md for the redistributed content covered by those terms. The videos, prompts, annotations, and metadata are all subject to the non-commercial research-only terms. Do not… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-Streaming-example-training-dataset.video-to-video1K<n<10K1 likes132 downloads3mo agoHugging Face26xrorrim /streaming_vlm_fix_1videon<1K0 likes128 downloads1y agoHugging Face27echodict /nemotron-3.5-asr-streaming-0.6b0 likes119 downloads4mo agoHugging Face28DataoceanAI /Chinese_Male_Speech_Synthesis_Corpus_Live_Streaming_for_Sales ID King-TTS-272 Duration 4.32 hours Language Chinese URL https://dataoceanai.com/datasets/tts/chinese-male-speech-synthesis-corpus-live-streaming-for-sales/ 7 likes94 downloads2y agoHugging Face29Jianguo-Huang11 /ego-streaming-bench-videosgatedvideo1K<n<10K0 likes92 downloads1mo agoHugging Face30JohnVitz /tsa-throughput-streaming-test3tabular10K<n<100K0 likes91 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.