Team Ai
Datasetpublic

changelinglab/cv-v1.0-segment

CommonVoice v1 Phone-Segment Alignments Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the charsiu/cv_ali release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09). Dataset summary lang train rows train hrs val rows val hrs test rows test hrs en 1,008,669 1,354.0 3,537 4.9 1,285 1.7 rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.

sourceHugging Facecc0-1.0updated 6mo agoView on Hugging Face
3likes1.9kdownloads
Dataset Card

CommonVoice v1 Phone-Segment Alignments

Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the `charsiu/cv_ali` release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09).

Dataset summary

langtrain rowstrain hrsval rowsval hrstest rowstest hrs
en1,008,6691,354.03,5374.91,2851.7
rw909,2521,115.500.000.0
ca863,4741,117.27,34510.24,3095.9
de539,246729.71,8252.5130.0
fr508,782605.11,9122.42420.3
be317,391360.61,6272.01430.2
es277,324343.91,0761.4500.1
it162,430208.33540.500.0
ba118,482120.84830.300.0
sw29,19439.34,7576.36970.9
TOTAL4,734,2445,994.522,91630.56,7399.1

Grand total: 4,763,899 utterances / 6,034.1 hours of aligned speech across train/val/test.

Schema

fieldtypenotes
utt_idstringCommonVoice clip stem (e.g. common_voice_en_12345)
audioAudio(sampling_rate=16000)mp3 bytes embedded in the parquet shards; resampled on decode
textstringsentence from CV {split}.tsv
phonesSequence[string]IPA phone labels from MFA
phone_startsSequence[float64]start time (seconds) of each phone
phone_endsSequence[float64]end time (seconds) of each phone
languagestringISO 639-1 code (ba, be, ca, ...)
speaker_idstringCV client_id (SHA hash)
durationfloat64last phone end time (seconds)
splitstringtrain, val, or test

Phone inventory is MFA's IPA output. Empty-label silence intervals from the source TextGrid are dropped.

Sources & attribution

  • —Audio & transcripts — Mozilla Common Voice Corpus 13.0, released 2023-03-09, under CC0 1.0.
  • —Phone-level alignments — `charsiu/cv_ali`, produced with MFA (Montreal Forced Aligner).

Citations

bibtex
@inproceedings{ardila-etal-2020-common,
    title = "{C}ommon {V}oice: A Massively-Multilingual Speech Corpus",
    author = "Ardila, Rosana  and
      Branson, Megan  and
      Davis, Kelly  and
      Kohler, Michael  and
      Meyer, Josh  and
      Henretty, Michael  and
      Morais, Reuben  and
      Saunders, Lindsay  and
      Tyers, Francis  and
      Weber, Gregor",
    booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference",
    month = may,
    year = "2020",
    address = "Marseille, France",
    publisher = "European Language Resources Association",
    url = "https://aclanthology.org/2020.lrec-1.520",
    pages = "4218--4222",
    language = "English",
    ISBN = "979-10-95546-34-4",
}

License

Released under CC0 1.0, matching the upstream Common Voice and Charsiu alignment licenses.