Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes6.9k downloads3y agoHugging Face02PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K49 likes5.2k downloads2y agoHugging Face03gilkeyio /librispeech-alignments Dataset Card for Librispeech Alignments Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here Dataset Details Dataset Description Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks. The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.audioautomatic-speech-recognition100K<n<1M21 likes3.3k downloads3y agoHugging Face04takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes2.4k downloads7mo agoHugging Face05AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes615 downloads5mo agoHugging Face06kyutai /interactivity-alignment-samples Audio Samples: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Audio samples accompanying the paper "Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models". Paper: arxiv.org Blog post: kyutai.org Models: 🤗 huggingface.co Overview This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/interactivity-alignment-samples.audio1K<n<10K9 likes445 downloads4mo agoHugging Face07QUD-Technologies /quran-alignment-benchmark Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.audioautomatic-speech-recognitionn<1K0 likes398 downloads6d agoHugging Face08heihei /hachimi-alignment Hachimi Alignment Dataset Accompanying dataset for "When Meaning Fades: Probing Acoustic Properties in Audio-Text Alignment" (ACL 2025). Paper and code: github.com/ngyygm/hachimi-alignment What are Hachimi Songs? Hachimi (哈基米) songs are Chinese internet parody songs that replace original meaningful lyrics with nonsense syllables ("ha-ji-mi") while preserving melody, rhythm, and vocal timbre. This creates a natural experiment for probing what audio-text alignment models… See the full description on the dataset page: https://huggingface.co/datasets/heihei/hachimi-alignment.audiofeature-extractionn<1K0 likes201 downloads6mo agoHugging Face09ErfanAShams /librispeech-alignments_clean100 librispeech-alignments_clean100 This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials. Cite: @inproceedings{panayotov2015librispeech, title={Librispeech: an ASR corpus based on public domain audio books}, author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev}, booktitle={ICASSP}, year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.audioautomatic-speech-recognition10K<n<100K1 likes188 downloads1y agoHugging Face10ivrit-ai /eval-forced-alignmentgated Hebrew Forced Alignment Evaluation Dataset Human-verified, word-level time-aligned Hebrew speech clips. To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was built. The system lets labelers fix the transcript and align each spoken word to the audio, down to 1ms precision (though annotators typically work at ~10ms granularity). The audio samples were gathered by randomly sampling from several of ivrit-ai's larger, published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.audioautomatic-speech-recognitionn<1K1 likes112 downloads17d agoHugging Face11Alignment-Lab-AI /librispeech-codec-22khzaudio10K<n<100K0 likes99 downloads10mo agoHugging Face12giangndm /audio-confidence-alignment Vietnamese Wav2Vec2 Feature & K-Means Tokenized Dataset This repository contains the structured speech features and tokenized cluster indices for the target pw733 and clean viVoice Vietnamese datasets, formatted as Parquet tables. 📊 Dataset Schema audio_uuid (string): Unique identifier of the audio file. text (string): Transcription text (empty for raw pw733 audio). features (list of list of float): Frame-level Wav2Vec2 embeddings ([Num_Frames, 768]). indices… See the full description on the dataset page: https://huggingface.co/datasets/giangndm/audio-confidence-alignment.text100K<n<1M0 likes55 downloads3mo agoHugging Face13Tuyentd /Post-Training_Answer_Style_Alignmentaudion<1K0 likes38 downloads1y agoHugging Face14The-Nature-of-Reality /THE-BLUEPRINT-FOR-AI-ALIGNMENTaudion<1K3 likes19 downloads2y agoHugging Face15Alignment-Lab-AI /podcast-1-test-preprocessedaudio1K<n<10K0 likes17 downloads2y agoHugging Face16nguyenvulebinh /libris-asr-alignmentaudion<1K0 likes14 downloads3y agoHugging Face17Alignment-Lab-AI /dogaudio10K<n<100K0 likes13 downloads3y agoHugging Face18sujalappa /sample-force-alignment-datasetaudion<1K0 likes10 downloads10mo agoHugging Face19sujalappa /new-forced-alignment-datasetaudion<1K0 likes9 downloads9mo agoHugging Face20Tuyentd /Conversational_Response_Style_Alignment_Resultaudion<1K0 likes8 downloads1y agoHugging Face21hungle161 /alignment_thai_error_testgatedaudio1K<n<10K0 likes6 downloads5mo agoHugging Face22AdoCleanCode /hifitts2_alignments_4750gatedaudio1K<n<10K0 likes5 downloads10mo agoHugging Face23AdoCleanCode /hifitts2_alignments_01gatedaudion<1K0 likes4 downloads10mo agoHugging Face24AdoCleanCode /hifitts2_alignments_0030gatedaudio10K<n<100K0 likes4 downloads10mo agoHugging Face25AdoCleanCode /hifitts2_alignments_6090gatedaudio10K<n<100K0 likes4 downloads10mo agoHugging Face26AdoCleanCode /hifitts2_alignments_70_75gatedaudio10K<n<100K0 likes4 downloads10mo agoHugging Face27AdoCleanCode /hifitts2_alignments_60100gatedaudio10K<n<100K0 likes4 downloads10mo agoHugging Face28AdoCleanCode /hifitts2_alignments_10_15gatedaudio1K<n<10K0 likes4 downloads10mo agoHugging Face29AdoCleanCode /hifitts2_alignments_05_10gatedaudio1K<n<10K0 likes4 downloads10mo agoHugging Face30AdoCleanCode /hifitts2_alignments_3060gatedaudio10K<n<100K0 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.