jml2026/multilingual-accent-speech-v2
Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed)… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech-v2.
Silencio Network: Multilingual Accent Speech Dataset (Sample)
<p align="left"> <img src="https://cdn-uploads.huggingface.co/production/uploads/69162b50b89e7abe20de4b5a/LWhs4p2lPFcyiVsP0tluu.png" width="40%"> </p>
 
Overview
Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which reduces legal risk for enterprise buyers. On top of that, the same community lets us scale quickly into hard-to-source languages and niches, so clients get both authenticity today and a credible path to large volumes tomorrow.
This dataset is a crowdsourced multilingual–accented English and non-English speech dataset designed for model training, benchmarking, and acoustic analysis. It emphasizes accent variation, short-form scripted prompts, and spontaneous free speech. All recordings were produced by contributors using their own devices, with Whisper-generated transcripts provided for every sample.
The dataset is structured for direct use in ASR, TTS, accent-classification, diarization-adjacent analysis, speech segmentation, and embedding evaluation.
Languages and Accents
This dataset covers five language–region pairs (to find out more about other combinations please reach out to us):
- English (China): English spoken with Mandarin-influenced accent
- English (Nigeria): Nigerian-accented English
- English (United States): Native American English speakers
- German (Germany): Native German speakers
- Spanish (Mexico): Native Mexican Spanish speakers
All recordings are stored as 48 kHz WAV files.
Usage
from datasets import load_dataset
# Load a specific language-region cohort
ds = load_dataset("SilencioNetwork/multilingual-accent-speech", "spanish_mexico")
# Access different script types using splits
for sample in ds['free_speech']:
audio = sample['audio'] # Audio data with sampling_rate and array
transcript = sample['transcript']
speaker_id = sample['speaker_id']
# Or access other splits
for sample in ds['keywords']:
# Process keywords samples
pass
for sample in ds['monologues']:
# Process monologues samples
passSpeech Types
Each sample belongs to one of three categories:
- free_speech: unscripted speech on a provided topic
- keywords: short isolated prompts containing specific phrases or terms
- monologues: longer scripted passages
These values appear in the field type_of_script.
Recording Conditions
All data is crowdsourced. Contributors record themselves using their available hardware and environment; conditions therefore vary naturally across microphones, devices, and noise profiles. No studio-grade normalisation or homogenisation is applied.
Transcription
Transcriptions are machine-generated using OpenAI Whisper, preserving its segmentation structure where applicable.
Dataset Statistics
This is a sample dataset organized by language-region, with each cohort split by script type. Durations are given in hours.
Spanish (Mexico)
English (China)
English (Nigeria)
English (United States)
German (Germany)
Overall Total
File Structure
Each config is organized by language-region, with splits for each script type:
spanish_mexico/
free_speech/
data/
audio_389928.wav
audio_390100.wav
... (25 files)
metadata.csv
keywords/
data/
audio_689765.wav
audio_706259.wav
... (6 files)
metadata.csv
monologues/
data/
audio_348730.wav
audio_348844.wav
... (25 files)
metadata.csv
english_china/
free_speech/
data/
audio_90018.wav
audio_90032.wav
... (25 files)
metadata.csv
keywords/
... (25 files)
monologues/
... (25 files)
... (3 more cohorts)Audio files are stored separately in AudioFolder format with metadata in CSV files.
Feature Schema
All configurations share the same feature structure:
- file_name: string (relative path to audio file)
- id: integer (unique identifier)
- speaker_id: string (hashed or anonymized speaker ID)
- gender: string (speaker gender)
- ethnicity: string (speaker ethnicity)
- occupation: string (occupation or profession)
- country_code: string (ISO 3166-1 alpha-2 code)
- birth_place: string (country or region of birth)
- mother_tongue: string (native language)
- dialect: string (regional dialect)
- yearofbirth: int (birth year, YYYY)
- yearsatbirth_place: int (years lived at birth place)
- languages_data: string (serialized language–proficiency data)
- os: string (recording operating system)
- device: string (recording device type)
- browser: string (browser used if web-based)
- duration: float (seconds) (audio length)
- emotions: string (emotion labels)
- language: string (primary language of the recording)
- location: string (recording location category)
- noise_sources: string (background noise labels)
- script_id: int (script template identifier)
- typeofscript: string {free_speech, keywords, monologues} (script category)
- script: string (text intended to be spoken)
- transcript: string (Whisper-generated transcription)
Licensing
Released under CC BY-NC 4.0. Commercial use is not permitted. Attribution to Silencio Network is required for any publication or derivative dataset.
Intended Use
Suitable for:
- accent-conditioned ASR training
- multilingual speech recognition
- TTS voicebank generation
- speaker embedding and similarity evaluation
- robustness benchmarking
- keyword-spotting models
- segmentation and VAD evaluation
Limitations
- Transcripts are automatically generated. Errors may be present.
- Crowdsourced device diversity introduces variable noise levels.
Contact
For questions, custom datasets, or commercial licensing inquiries, please visit our website.
Citation
@dataset{silencio_network_speech_2025,
title = {Silencio Network Multilingual Accent Speech Corpus},
author = {Silencio Network},
year = {2025},
license = {CC BY-NC 4.0}
}