Team Ai
Datasetpublic

jml2026/multilingual-accent-speech-v2

Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed)… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech-v2.

sourceHugging Facecc-by-nc-4.0updated 8mo agoView on Hugging Face
0likes46downloads
Dataset Card

Silencio Network: Multilingual Accent Speech Dataset (Sample)

<p align="left"> <img src="https://cdn-uploads.huggingface.co/production/uploads/69162b50b89e7abe20de4b5a/LWhs4p2lPFcyiVsP0tluu.png" width="40%"> </p>

![Website](https://www.silencioai.com) ![Contact](https://www.silencioai.com/contact)

Overview

Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which reduces legal risk for enterprise buyers. On top of that, the same community lets us scale quickly into hard-to-source languages and niches, so clients get both authenticity today and a credible path to large volumes tomorrow.

This dataset is a crowdsourced multilingual–accented English and non-English speech dataset designed for model training, benchmarking, and acoustic analysis. It emphasizes accent variation, short-form scripted prompts, and spontaneous free speech. All recordings were produced by contributors using their own devices, with Whisper-generated transcripts provided for every sample.

The dataset is structured for direct use in ASR, TTS, accent-classification, diarization-adjacent analysis, speech segmentation, and embedding evaluation.

Languages and Accents

This dataset covers five language–region pairs (to find out more about other combinations please reach out to us):

  • —English (China): English spoken with Mandarin-influenced accent
  • —English (Nigeria): Nigerian-accented English
  • —English (United States): Native American English speakers
  • —German (Germany): Native German speakers
  • —Spanish (Mexico): Native Mexican Spanish speakers

All recordings are stored as 48 kHz WAV files.

Usage

python
from datasets import load_dataset

# Load a specific language-region cohort
ds = load_dataset("SilencioNetwork/multilingual-accent-speech", "spanish_mexico")

# Access different script types using splits
for sample in ds['free_speech']:
    audio = sample['audio']  # Audio data with sampling_rate and array
    transcript = sample['transcript']
    speaker_id = sample['speaker_id']

# Or access other splits
for sample in ds['keywords']:
    # Process keywords samples
    pass

for sample in ds['monologues']:
    # Process monologues samples
    pass

Speech Types

Each sample belongs to one of three categories:

  • —free_speech: unscripted speech on a provided topic
  • —keywords: short isolated prompts containing specific phrases or terms
  • —monologues: longer scripted passages

These values appear in the field type_of_script.

Recording Conditions

All data is crowdsourced. Contributors record themselves using their available hardware and environment; conditions therefore vary naturally across microphones, devices, and noise profiles. No studio-grade normalisation or homogenisation is applied.

Transcription

Transcriptions are machine-generated using OpenAI Whisper, preserving its segmentation structure where applicable.

Dataset Statistics

This is a sample dataset organized by language-region, with each cohort split by script type. Durations are given in hours.

Spanish (Mexico)

SplitRecordingsSpeakersDuration (hrs)
free_speech2550.27
keywords620.05
monologues2570.45
Subtotal5690.77

English (China)

SplitRecordingsSpeakersDuration (hrs)
free_speech25130.33
keywords2560.19
monologues25100.44
Subtotal75180.96

English (Nigeria)

SplitRecordingsSpeakersDuration (hrs)
free_speech25230.32
keywords25230.16
monologues25210.53
Subtotal75461.01

English (United States)

SplitRecordingsSpeakersDuration (hrs)
free_speech25180.3
keywords25140.18
monologues25130.32
Subtotal75310.80

German (Germany)

SplitRecordingsSpeakersDuration (hrs)
free_speech25160.25
keywords25150.16
monologues25140.32
Subtotal75270.73

Overall Total

MetricValue
Total Recordings356
Total Speakers161
Total Duration4.29 hrs
Configs5
Splits per Config3

File Structure

Each config is organized by language-region, with splits for each script type:

spanish_mexico/
    free_speech/
        data/
            audio_389928.wav
            audio_390100.wav
            ... (25 files)
        metadata.csv
    keywords/
        data/
            audio_689765.wav
            audio_706259.wav
            ... (6 files)
        metadata.csv
    monologues/
        data/
            audio_348730.wav
            audio_348844.wav
            ... (25 files)
        metadata.csv
english_china/
    free_speech/
        data/
            audio_90018.wav
            audio_90032.wav
            ... (25 files)
        metadata.csv
    keywords/
        ... (25 files)
    monologues/
        ... (25 files)
... (3 more cohorts)

Audio files are stored separately in AudioFolder format with metadata in CSV files.

Feature Schema

All configurations share the same feature structure:

  • —file_name: string (relative path to audio file)
  • —id: integer (unique identifier)
  • —speaker_id: string (hashed or anonymized speaker ID)
  • —gender: string (speaker gender)
  • —ethnicity: string (speaker ethnicity)
  • —occupation: string (occupation or profession)
  • —country_code: string (ISO 3166-1 alpha-2 code)
  • —birth_place: string (country or region of birth)
  • —mother_tongue: string (native language)
  • —dialect: string (regional dialect)
  • —yearofbirth: int (birth year, YYYY)
  • —yearsatbirth_place: int (years lived at birth place)
  • —languages_data: string (serialized language–proficiency data)
  • —os: string (recording operating system)
  • —device: string (recording device type)
  • —browser: string (browser used if web-based)
  • —duration: float (seconds) (audio length)
  • —emotions: string (emotion labels)
  • —language: string (primary language of the recording)
  • —location: string (recording location category)
  • —noise_sources: string (background noise labels)
  • —script_id: int (script template identifier)
  • —typeofscript: string {free_speech, keywords, monologues} (script category)
  • —script: string (text intended to be spoken)
  • —transcript: string (Whisper-generated transcription)

Licensing

Released under CC BY-NC 4.0. Commercial use is not permitted. Attribution to Silencio Network is required for any publication or derivative dataset.

Intended Use

Suitable for:

  • —accent-conditioned ASR training
  • —multilingual speech recognition
  • —TTS voicebank generation
  • —speaker embedding and similarity evaluation
  • —robustness benchmarking
  • —keyword-spotting models
  • —segmentation and VAD evaluation

Limitations

  • —Transcripts are automatically generated. Errors may be present.
  • —Crowdsourced device diversity introduces variable noise levels.

Contact

For questions, custom datasets, or commercial licensing inquiries, please visit our website.

Citation

@dataset{silencio_network_speech_2025,
    title        = {Silencio Network Multilingual Accent Speech Corpus},
    author       = {Silencio Network},
    year         = {2025},
    license      = {CC BY-NC 4.0}
}