Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes3.5k downloads6mo agoHugging Face02OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition98 likes2.2k downloads6mo agoHugging Face03chikingsley /l2-arctic-release-v5.0 L2-ARCTIC v5.0 L2-ARCTIC is a non-native English speech corpus intended for research in pronunciation assessment, accent conversion, voice conversion, and mispronunciation detection. This Hub dataset mirrors the v5.0 release as distributed locally: speaker-level zip archives, the suitcase corpus archive, prompts, the original README, and the license. Summary 24 non-native English speakers 26,867 utterances 27.1 hours of scripted speech Manual phone-level annotations for… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-release-v5.0.automatic-speech-recognition10K<n<100K0 likes798 downloads7mo agoHugging Face04Archime /french_tv_media_dataset_2026 Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus Résumé (Abstract) Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.audioautomatic-speech-recognition10K<n<100K5 likes362 downloads8mo agoHugging Face05chikingsley /l2-arctic-manual-v5.0-16k l2-arctic-manual-v5.0-16k This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only the manually annotated material and converts the audio to 16 kHz mono FLAC. It is designed to plug into the current peacock-asr training code, which can consume a Hugging Face dataset with audio plus phonemes. Included splits train: 1800 rows, 1.84 hours validation: 899 rows, 0.94 hours test: 900 rows, 0.88 hours suitcase: 22 rows, 0.44 hours The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.audioautomatic-speech-recognition1K<n<10K0 likes203 downloads7mo agoHugging Face06arcada-labs /conversation-bench Conversation Bench 75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.audioautomatic-speech-recognitionn<1K8 likes167 downloads7mo agoHugging Face07arcada-labs /grocery-bench Grocery Bench 30-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a grocery ordering assistant. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a grocery ordering assistant helping a customer build, modify, and finalize an order. The conversation is designed around 15 difficulty enhancements that… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/grocery-bench.audioautomatic-speech-recognitionn<1K2 likes150 downloads7mo agoHugging Face08outlawmold /sinhala-tts-dataset-archive-20260429-082457 Sinhala TTS Dataset Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Stats Metric Value Utterances 218 Train 208 Val 10 Hours 0.51 Mean duration 8.5s Sample rate 22050 Hz Pipeline Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 -> Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB) Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.audiotext-to-speechn<1K0 likes135 downloads6mo agoHugging Face09KeisukeMiyamoto /nhk-archive-audio-30sgated NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.audioautomatic-speech-recognition100K<n<1M0 likes107 downloads15d agoHugging Face10arcada-labs /appointment-bench Appointment Bench 25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.audioautomatic-speech-recognitionn<1K5 likes100 downloads7mo agoHugging Face11arcada-labs /product-bench Product Bench 31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.audioautomatic-speech-recognitionn<1K2 likes91 downloads7mo agoHugging Face12arcada-labs /assistant-bench Assistant Bench 31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.audioautomatic-speech-recognitionn<1K2 likes76 downloads7mo agoHugging Face13arcada-labs /event-bench Event Bench 29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.audioautomatic-speech-recognitionn<1K2 likes49 downloads7mo agoHugging Face14Zeldeo /transatlantic-voice-archive_distille Distillation brucemacd/transatlantic-voice-archive Dataset ASR distillé via Cohere Transcribe. Source : brucemacd/transatlantic-voice-archive Modèle ASR : cohere-transcribe Langue ASR : en Exemples : 1427 (dataset source intégral) Colonnes : audio — clip audio (16 kHz) transcription_base — référence brute du dataset source transcription_cohere — hypothèse Cohere brute langue_accent — langue / accent détecté wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.audioautomatic-speech-recognition1K<n<10K0 likes41 downloads2mo agoHugging Face15archivartaunik /be-sidon-restored-sample-100 be-sidon-restored-sample-100 Набор прыкладаў беларускай мовы з Common Voice (validated), апрацаваны мадэллю аднаўлення (denoise). Файлы audio/original/ — арыгінальныя кліпы (як у Common Voice). audio/restored/ — адноўленыя WAV (48 кГц). metadata.csv — палеткі: audio, original_audio, sentence, speaker. Дата стварэння: 2025-09-25. automatic-speech-recognitionn<1K0 likes26 downloads1y agoHugging Face16arcada-labs /audio-agent-bench-suite Audio Agent Bench Suite A suite of six multi-turn, multi-domain spoken conversational benchmarks for evaluating voice AI and audio agent systems. Each sub-dataset targets a distinct real-world deployment domain, together covering the core capabilities required of production audio agents: instruction following, knowledge-base grounding, tool/function-call accuracy, long-range conversational memory, and state tracking. Sub-datasets Dataset Domain Turns HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/audio-agent-bench-suite.automatic-speech-recognition0 likes23 downloads5mo agoHugging Face17archivartaunik /ZhygamontCmok ZhygamontCmok Raw WAV files uploaded directly to the dataset repository. Structure audio/ — merged WAV files audioautomatic-speech-recognitionn<1K0 likes18 downloads6mo agoHugging Face18brucemacd /transatlantic-voice-archive Transatlantic Voice Archive Speech clips with aligned transcripts in the transatlantic (Mid-Atlantic) accent — the clipped, semi-British delivery of 1930s–1960s American newsreel announcers. Built from public-domain Universal Newsreels (1929–1967) on the Internet Archive, intended for finetuning TTS models on the accent. Dataset statistics Clips 1427 Total audio 2.04 h (122.2 min) Average clip 5.14 s Sample rate 22050 Hz, mono WAV Source reels… See the full description on the dataset page: https://huggingface.co/datasets/brucemacd/transatlantic-voice-archive.audiotext-to-speech1K<n<10K0 likes16 downloads4mo agoHugging Face19Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05 This dataset contains lossless FLAC chunks derived from 50 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 1,536 Chunk audio duration: 5.773 hours Source transcript rows represented: 4,326 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 22 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05.audiotext-to-speech1K<n<10K0 likes15 downloads2mo agoHugging Face20Prompthumanizer /jain_architecturegatedtoken-classificationn>1T0 likes12 downloads1y agoHugging Face21Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01 This dataset contains lossless FLAC chunks derived from 45 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 1,034 Chunk audio duration: 4.411 hours Source transcript rows represented: 3,041 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 9 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01.audiotext-to-speech1K<n<10K0 likes12 downloads2mo agoHugging Face22Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02 This dataset contains lossless FLAC chunks derived from 50 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 1,097 Chunk audio duration: 5.143 hours Source transcript rows represented: 3,875 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 19 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02.audiotext-to-speech1K<n<10K0 likes12 downloads2mo agoHugging Face23Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06 This dataset contains lossless FLAC chunks derived from 50 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 2,116 Chunk audio duration: 6.624 hours Source transcript rows represented: 5,601 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 56 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06.audiotext-to-speech1K<n<10K0 likes12 downloads2mo agoHugging Face24Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03 This dataset contains lossless FLAC chunks derived from 50 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 650 Chunk audio duration: 5.893 hours Source transcript rows represented: 4,213 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 1 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03.audiotext-to-speechn<1K0 likes11 downloads2mo agoHugging Face25Kppwdfgu1 /gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04gated gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04 This dataset contains lossless FLAC chunks derived from 50 Nigerian-language, Nigerian English, and Nigerian Pidgin recordings. Access requests require manual approval by the repository owner. Contents Chunks: 599 Chunk audio duration: 4.567 hours Source transcript rows represented: 2,961 Standalone audio-tag chunks: 0 Non-music tag rows attached to nearest same-speaker speech: 0 Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04.audiotext-to-speechn<1K0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.