Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for 8 language subsets, with 10,100 examples selected from Common Voice 17.0. Each example supplies a speaker reference and an independently selected target text, with human recordings as WER/CER and speaker-similarity anchors when the corresponding audio is available. Corpus WER and the WavLM-FT evaluation follow seed-tts-eval. Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes1.8k downloads4d agoHugging Face02sophia8888 /clipquill-asr-benchmark Measuring whisper-tiny vs whisper-base in a browser tab Word error rate, wall-clock timing, transfer size and peak memory for two quantised Whisper tiers running entirely client-side in a real Chrome window, with the scripts that produced every number. If you are building an in-browser transcription page, the two results worth knowing before you pick a model tier: On clean synthetic audio the two tiers tie. If that is all you test, you will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.tabularautomatic-speech-recognitionn<1K0 likes152 downloads22d agoHugging Face03SaarAI /asr-benchmark-outputsgated SaarAI ASR Benchmark Outputs Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets. files: 570 utterances: 4615160 languages: 7 models: 50 Layout data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl index.jsonl # one record per file (language, split, model, rows, sha256, ...) index.csv Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tabularautomatic-speech-recognition1M<n<10M1 likes113 downloads5d agoHugging Face04themechanism /script-fidelity-benchmark Script fidelity benchmark Anonymous supplement for the paper "Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark." Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters that belong to the expected target script. WER measures word edits, while SFR checks whether the output is written in the target orthography. Related resources: PyPI package: https://pypi.org/project/script-fidelity/ Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes81 downloads5mo agoHugging Face05sumanpaudel1997 /nepali-asr-benchmark Nepali ASR Benchmark Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR checkpoints evaluated on three independent test sets. Released alongside the paper Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition. Contents Field Type Description utterance_id string stable identifier {test_set}-{index} reference string NFC-normalised gold transcription (Devanagari) hypothesis string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes44 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.