Team Ai
Datasetpublic

thomasgauthier/common-voice-scripted-speech-quebec

common-voice-scripted-speech-quebec This dataset is a filtered subset of the Mozilla Common Voice Scripted Speech 25.0 - French dataset. It exclusively contains audio clips from speakers with Canadian and Québécois accents. Dataset Summary Language: French (fr) Total Clips: 25,198 Total Duration: 36.83 hours (132,578.64 seconds) License: CC-0 Filtering Criteria This subset was generated by extracting rows from the original… See the full description on the dataset page: https://huggingface.co/datasets/thomasgauthier/common-voice-scripted-speech-quebec.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes100downloads
Dataset Card

common-voice-scripted-speech-quebec

This dataset is a filtered subset of the Mozilla Common Voice *Scripted Speech* 25.0 - French dataset. It exclusively contains audio clips from speakers with Canadian and Québécois accents.

Dataset Summary

  • —Language: French (fr)
  • —Total Clips: 25,198
  • —Total Duration: 36.83 hours (132,578.64 seconds)
  • —License: CC-0

Filtering Criteria

This subset was generated by extracting rows from the original cv-corpus-25.0-2026-03-09 dataset that match the following criteria:

  • —variant is exactly "Français d'Amérique du Nord"
  • —OR accents matches any self-reported Canadian/Québécois identifier (e.g., "Français du Canada", "Québécois", "Français du Québec|Saguenay").

Data Structure

The dataset is formatted as a Hugging Face Dataset with the audio cast to the Audio() feature type.

FeatureTypeDescription
audioAudioThe loaded audio dictionary (bytes/path) and sampling rate
transcriptionstringThe text transcription of the audio
speaker_idstringHashed UUID of the contributor
sentence_idstringUnique identifier for the text prompt
sentence_domainstringDomain(s) the sentence belongs to
up_votesint64Number of users who validated the clip
down_votesint64Number of users who rejected the clip
agestringSelf-declared age bracket
genderstringSelf-declared gender
accentsstringSelf-declared accent
variantstringSelf-declared language variant
localestringOriginal locale code (fr)
segmentstringCustom dataset segment identifier
splitstringOriginal dataset split (e.g., train, dev, test, validated)

Source and Documentation

For comprehensive information regarding the original data collection methodology, text corpus sources, demographic definitions, and broader project details, please refer to the official datasheet:

[https://mozilladatacollective.com/datasets/cmn5zugst00w3nv07upovf2bg](https://mozilladatacollective.com/datasets/cmn5zugst00w3nv07upovf2bg)