Borrison/hk-legicost
HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level. Paper: arXiv:2306.11252 Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur Dataset Description The raw… See the full description on the dataset page: https://huggingface.co/datasets/Borrison/hk-legicost.
HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation
HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level.
Paper: arXiv:2306.11252
Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur
Dataset Description
The raw data consists of recordings of the Hong Kong Legislative Council's (LegCo) meetings, which focuses mainly on government policy-related inquiries and their corresponding responses, alongside discussions and debates on motions and resolutions.
Cantonese is a language whose written forms sometimes deviate from its spoken form. Specifically, its written form is often altered to appear closer to the Mandarin written form. We refer to this written form as standard Chinese and say these transcripts are "non-verbatim" to reflect the discrepancy between what is spoken and written. This makes the corpus particularly suitable for speech translation research when there are significant differences between the spoken and written forms of the source language.
Speech-to-text and bitext alignment were performed with automatic approaches, with a filtering step for discarding alignment results whose confidence score is below a threshold determined empirically. As the transcripts are not verbatim, there may exist utterances whose transcript is not perfectly aligned.
Splits
The splits are created using the automatic tool nachos such that:
- dev_asr splits are speaker-disjoint with respect to the train split
- dev_mt splits are document-disjoint with respect to the train split
- test split is both speaker and document disjoint with respect to the train split
The splits match the statistics reported in the original paper. All audio files, transcripts, translations, and metadata are included in full.
Data Format
Each split is stored as a Parquet file (data/<split>.parquet) with the following columns:
CSV versions are also available at data/<split>.csv (tab-separated).
Audio Format
All audio files are mono WAV files at 16kHz sample rate, 16-bit PCM encoding. They are stored in data/audio/.
Metadata
<split>.utt.metadata— Mapping from hashed audio filename to original filename (which may include topic/speaker information)<split>.spk.metadata— Mapping from hashed speaker ID to true speaker identity
Usage
from datasets import load_dataset
# Load the full dataset (audio decoded automatically)
dataset = load_dataset("Borrison/hk-legicost")
# Load a specific split
test = load_dataset("Borrison/hk-legicost", split="test")
# Access an example
example = test[0]
print(example["transcript"]) # Traditional Chinese
print(example["translation"]) # English
print(example["audio"]) # Decoded audio dict with 'array', 'sampling_rate', 'path'To load just the metadata without audio decoding:
dataset = load_dataset("Borrison/hk-legicost", streaming=True)Citation
@misc{xiao2023hklegicost,
title={HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation},
author={Cihan Xiao and Henry Li Xinyuan and Jinyi Yang and Dongji Gao and Matthew Wiesner and Kevin Duh and Sanjeev Khudanpur},
year={2023},
eprint={2306.11252},
archivePrefix={arXiv},
primaryClass={cs.CL}
}License
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Acknowledgements
The data is sourced from recordings of the Hong Kong Legislative Council meetings, which are publicly available on the LegCo website.
