vikkyblacq/kare-codeswitch-samples
Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.
Kare — Code-Switching Illustrative Samples
Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge.
What this is — and isn't
- Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin with embedded English clinical terms — the exact phenomenon Kare's ASR benchmark exists to measure. Audio was synthesized with Sahara (Intron) TTS, the speech API this challenge provides.
- Isn't the benchmark corpus itself. Kare's actual benchmark evaluates five ASR models against Intron's own
AfriSwitchCareandAfriSwitchdatasets — real, professionally-produced datasets gated on Hugging Face (contact-approval and manual-approval respectively). Those aren't ours to redistribute; this repo links to the source instead: - https://huggingface.co/datasets/intronhealth/AfriSwitchCare
- https://huggingface.co/datasets/intronhealth/AfriSwitch
This dataset is a small, original, fully-owned set built to give reviewers something to actually listen to without going through that gate. No human speaker's voice or identity is involved — every clip is machine-synthesized from text we wrote — so there is no consent or de-identification question to begin with.
Contents
8 clips, ~35 seconds total, 2 per language (Yoruba, Hausa, Igbo, Nigerian Pidgin), each pairing everyday matrix-language grammar with embedded English clinical vocabulary (symptom names, drug names). See metadata.csv for the full table (file_name, language_code, language, duration_s, text, english_gloss, voice).
English glosses for all eight are in metadata.csv.
A note on accuracy
The Pidgin and Yoruba lines were written and checked by the team directly; Hausa and Igbo constructions were written in-house for this demo and are common, everyday phrasing, but have not been reviewed by a native or professional linguist. If something reads oddly to a fluent speaker, treat it as a demo-quality artifact, not a claim of linguistic authority.
Provenance & reuse
- Text: original, written by the Kare team for this submission — free to reuse or adapt.
- Audio: synthesized via Sahara (Intron) TTS under Kare's access to that API for the Sahara CodeSwitch Africa Challenge; shared here for challenge review. If you want to reuse the audio itself outside that context, check with Intron on Sahara TTS output terms first.
- Loading: standard Hugging Face
audiofolderlayout —load_dataset("vikkyblacq/kare-codeswitch-samples")picks upmetadata.csvautomatically.
Built for the Sahara CodeSwitch Africa Challenge. Kare's full benchmark, code and submission materials: https://github.com/Vikky-Ajayi/Kare
