Kennethdot/Ghana_English-Twi_Code-switching_Speech
Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.
Dataset Card for KasaSpeech
Dataset Summary
KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios
With over 95 hours of manually transcribed speech, KasaSpeech establishes a new benchmark and gold-standard corpus for English–Twi code-switching speech recognition and text-to-speech research. It is designed to support the development, evaluation, and comparison of ASR systems, speech representation models, and multilingual speech technologies for English–Twi.
Supported Tasks
KasaSpeech is suitable for:
- Automatic Speech Recognition (ASR) for Code-switching speech
- Text-To-Speech (TTS)
- Multilingual speech modeling
- Speech representation learning
- Speech foundation model fine-tuning and evaluation
- African language speech technology research
Dataset Structure
Data Splits
Data Fields
Each example contains the following fields:
Example
from datasets import load_dataset, Audio
dataset = load_dataset(
"Kennethdot/Ghana_English-Twi_Code_switching_ASR",
split="train"
)
dataset = dataset.cast_column(
"audio",
Audio(sampling_rate=16000)
)
sample = dataset[0]
print(sample["transcript"])Example transcript:
Me phone no a-crack-i, henfa na mɛtumi a-fix-i screen no?Dataset Creation
Collection Process
Speech recordings were voluntarily contributed by participants using a custom data collection platform. Speakers were presented with prompts designed to encourage natural English–Twi code-switching while covering a broad range of everyday topics and communication scenarios.
Annotation Process
All recordings were manually transcribed following standardized annotation guidelines developed for English–Twi code-switched speech. Multiple quality assurance steps were performed to improve transcription consistency and remove corrupted or invalid recordings.
Speaker Information
The dataset includes recordings from speakers spanning multiple age groups and genders. Speaker identities have been anonymized using unique identifiers.
Dataset Characteristics
- Total recordings: 54,855
- Total duration: 95.58 hours
- Languages: English, Twi, and English–Twi code-switching
- Sampling rate: 48 kHz (can be resampled to 16 kHz for model training)
- Recording style: Prompted, natural code-switched speech
- Transcriptions: Human-annotated
Limitations
- Demographic representation may not be perfectly balanced across speaker groups.
- Recording conditions vary across devices and environments.
- The dataset primarily reflects Ghanaian English–Twi code-switching and may not generalize to all Akan dialects or other multilingual contexts.
- Although carefully curated, minor transcription inconsistencies may remain.
Citation
If you use KasaSpeech in your work, please cite:
@dataset{kasaspeech2026,
title={KasaSpeech: A Large-Scale English--Twi Code-Switching Speech Dataset},
author={Dotse, Kenneth},
year={2026},
url={https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code_switching_ASR}
}Contact
For questions, bug reports, or collaboration opportunities, please open a discussion on the Hugging Face dataset page. Contributions, feedback, and research collaborations are welcome.
