Team Ai
Datasetpublic

Borrison/hk-legicost

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level. Paper: arXiv:2306.11252 Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur Dataset Description The raw… See the full description on the dataset page: https://huggingface.co/datasets/Borrison/hk-legicost.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
3likes1kdownloads
Dataset Card

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation

HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level.

Paper: arXiv:2306.11252

Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur

Dataset Description

The raw data consists of recordings of the Hong Kong Legislative Council's (LegCo) meetings, which focuses mainly on government policy-related inquiries and their corresponding responses, alongside discussions and debates on motions and resolutions.

Cantonese is a language whose written forms sometimes deviate from its spoken form. Specifically, its written form is often altered to appear closer to the Mandarin written form. We refer to this written form as standard Chinese and say these transcripts are "non-verbatim" to reflect the discrepancy between what is spoken and written. This makes the corpus particularly suitable for speech translation research when there are significant differences between the spoken and written forms of the source language.

Speech-to-text and bitext alignment were performed with automatic approaches, with a filtering step for discarding alignment results whose confidence score is below a threshold determined empirically. As the transcripts are not verbatim, there may exist utterances whose transcript is not perfectly aligned.

Splits

SplitUtterancesHoursDescription
train142,432518.6Training set
devasr03,33812.5ASR dev (speaker-disjoint)
devasr13,33812.4ASR dev (speaker-disjoint)
devasr24,46916.9ASR dev (speaker-disjoint)
devmt03,19411.6MT dev (document-disjoint)
devmt13,19411.9MT dev (document-disjoint)
devmt24,38416.0MT dev (document-disjoint)
test2,5129.5Test (speaker + document disjoint)
Total166,861609.4

The splits are created using the automatic tool nachos such that:

  • —dev_asr splits are speaker-disjoint with respect to the train split
  • —dev_mt splits are document-disjoint with respect to the train split
  • —test split is both speaker and document disjoint with respect to the train split

The splits match the statistics reported in the original paper. All audio files, transcripts, translations, and metadata are included in full.

Data Format

Each split is stored as a Parquet file (data/<split>.parquet) with the following columns:

ColumnTypeDescription
audiostringRelative path to WAV file in data/audio/
speakerstringHashed speaker ID
start_timefloat32Start time of the utterance in seconds
stop_timefloat32Stop time of the utterance in seconds
transcriptstringTraditional Chinese transcript
translationstringEnglish translation

CSV versions are also available at data/<split>.csv (tab-separated).

Audio Format

All audio files are mono WAV files at 16kHz sample rate, 16-bit PCM encoding. They are stored in data/audio/.

Metadata

  • —<split>.utt.metadata — Mapping from hashed audio filename to original filename (which may include topic/speaker information)
  • —<split>.spk.metadata — Mapping from hashed speaker ID to true speaker identity

Usage

python
from datasets import load_dataset

# Load the full dataset (audio decoded automatically)
dataset = load_dataset("Borrison/hk-legicost")

# Load a specific split
test = load_dataset("Borrison/hk-legicost", split="test")

# Access an example
example = test[0]
print(example["transcript"])   # Traditional Chinese
print(example["translation"])  # English
print(example["audio"])        # Decoded audio dict with 'array', 'sampling_rate', 'path'

To load just the metadata without audio decoding:

python
dataset = load_dataset("Borrison/hk-legicost", streaming=True)

Citation

bibtex
@misc{xiao2023hklegicost,
      title={HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation},
      author={Cihan Xiao and Henry Li Xinyuan and Jinyi Yang and Dongji Gao and Matthew Wiesner and Kevin Duh and Sanjeev Khudanpur},
      year={2023},
      eprint={2306.11252},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

License

This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

Acknowledgements

The data is sourced from recordings of the Hong Kong Legislative Council meetings, which are publicly available on the LegCo website.