Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Parakeet-Inc /joyo-kanji-yomi-benchmark-parakeet 日本語 | English 常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet) 常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。 このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。 このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。 概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.texttext-to-speech10K<n<100K5 likes334 downloads2mo agoHugging Face02ketav /parakeet-hindi-asr Parakeet Hindi-English Bilingual ASR Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition. Quick Start # Download pip install huggingface_hub huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr # Install dependencies pip install nemo_toolkit[asr] bitsandbytes sentencepiece # Train (after updating paths in config) cd parakeet-hindi-asr/scripts python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.textautomatic-speech-recognition100K<n<1M0 likes67 downloads9mo agoHugging Face03Parakeet-Inc /J-HARD-TTS-Eval J-HARD-TTS-Eval [!NOTE] For full documentation, detailed benchmark results, and methodology, please refer to the GitHub Repository. Overview J-HARD-TTS-Eval is a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models. It focuses on specific failure modes such as stability in short sequences, repetition handling, and context completion. Usage You can easily load the dataset using the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/J-HARD-TTS-Eval.audiotext-to-speechn<1K6 likes56 downloads8mo agoHugging Face040x3 /joyo-kanji-yomi-benchmark-parakeet 日本語 | English 常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet) 常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。 このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。 このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。 概要… See the full description on the dataset page: https://huggingface.co/datasets/0x3/joyo-kanji-yomi-benchmark-parakeet.texttext-to-speech10K<n<100K0 likes45 downloads1mo agoHugging Face05NeurologyAI /neuro-parakeet-foodgated neuro-whisper-v1 Dataset Description This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains. Data Generation Voice Data: Synthetically generated using Resemble AI Chatterbox TTS Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.audioautomatic-speech-recognition10K<n<100K4 likes29 downloads9mo agoHugging Face06kapilkarda /parakeet-indic-audioaudio0 likes25 downloads6mo agoHugging Face07TieIncred /parakeet-tdt-blind-spots Blind Spots of nvidia/parakeet-tdt-0.6b-v2 This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data. Model Under Test Property Value Model nvidia/parakeet-tdt-0.6b-v2 Parameters 600M Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.audioautomatic-speech-recognitionn<1K0 likes24 downloads7mo agoHugging Face08otoha-project /moe-speech-plus-cache-parakeet-v1-auditgated MoeSpeechPlus Parakeet cache v1 — aggregate audit This aggregate-only audit describes otoha-project/moe-speech-plus-cache-parakeet-v1 at revision fa80a0e57e4af06df434130f66a2792207f00c45. It excludes speaker IDs, utterance IDs, paths, text, tensors, pickle data, and cache files. The two tar parts form a continuous stream and terminate correctly. The repository also contains a 31,232-file expanded tree; one sampled expanded file was byte-identical to its tar member. The cache… See the full description on the dataset page: https://huggingface.co/datasets/otoha-project/moe-speech-plus-cache-parakeet-v1-audit.tabularn<1K0 likes22 downloads14d agoHugging Face09Trelis /eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 11.34% 3.63% Source Data Evaluation Dataset: Trelis/medical-terms-2025 Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926.audion<1K1 likes21 downloads6mo agoHugging Face10toth235a /parakeet-whisper-divergence Parakeet vs Whisper Divergence Samples This dataset contains 40 audio samples where Parakeet TDT v3 and Whisper (Granary pseudolabels) show significant divergence. Purpose Investigate why Whisper outputs blank or very short transcriptions on French VoxPopuli audio. Issues Found whisper_short: Whisper outputs < 2 chars/sec while Parakeet outputs > 5 chars/sec high_wer: WER > 80% between Whisper and Parakeet transcriptions Common Whisper Hallucinations… See the full description on the dataset page: https://huggingface.co/datasets/toth235a/parakeet-whisper-divergence.0 likes20 downloads9mo agoHugging Face11rdsm /parakeet-stt-redone parakeet-stt-redone What this is 108,276 raw→clean transcript pairs sourced from aldigobbler/stt-correction, re-labeled using GLM-5.1-FP8 as the teacher model with our production cleanup prompt. How it differs from the source dataset aldigobbler/stt-correction this dataset Target Verbatim transcript restoration (lowercase, no punctuation, fillers kept/restored) Polished readable text — punctuated, paragraphed, fillers selectively removed… See the full description on the dataset page: https://huggingface.co/datasets/rdsm/parakeet-stt-redone.texttext-generation100K<n<1M0 likes13 downloads4mo agoHugging Face12Trelis /eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 37.59% 20.64% Source Data Evaluation Dataset: Trelis/eka-hard Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920.audion<1K0 likes11 downloads6mo agoHugging Face13Trelis /eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 15.94% 10.13% Source Data Evaluation Dataset: Trelis/multimed-hard Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930.audion<1K0 likes11 downloads6mo agoHugging Face14gab1k /mmm_project_parakeetimagen<1K0 likes5 downloads10mo agoHugging Face15otoha-project /moe-speech-plus-cache-parakeet-v1gated moe-speech-plus cache (parakeet transcription・cache_version=1) otoha m5_f5_trainer の Path B' precache 生成物。 内訳 cache_version: 1 transcription_source: parakeet n_samples: 354,110 speakers: 424 (moe-speech-plus 全 473 UUID から hold-out 48 話者除外・fair 分布) hold-out UUID: 48 (GAP-11 で確定・commit 3f3a703) 使い方 from huggingface_hub import snapshot_download cache_dir = snapshot_download( repo_id="otoha-project/moe-speech-plus-cache-parakeet-v1"… See the full description on the dataset page: https://huggingface.co/datasets/otoha-project/moe-speech-plus-cache-parakeet-v1.0 likes5 downloads3mo agoHugging Face16gab1k /mmm_project_parakeet_intermdataimagen<1K0 likes4 downloads10mo agoHugging Face17ketav /hinglish-parakeet-tarredgatedaudio1M<n<10M0 likes4 downloads7mo agoHugging Face18kapilkarda /parakeet-indic-datatext10M<n<100M0 likes3 downloads6mo agoHugging Face19vambassa /parakeet_inferencetext1K<n<10K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.