bemba
Datasets
All datasets matching “bemba”BEMBA_big_cbemba-speech-csikasote
Description
This is speech dataset of Bemba language. This dataset was acquired from (BembaSpeech)[https://github.com/csikasote/BembaSpeech/tree/master].
BembaSpeech is the speech recognition corpus in Bemba [1].
Dataset Structure
DatasetDict({
train: Dataset({
features: ['audio', 'sentence'],
num_rows: 12421
})
dev: Dataset({
features: ['audio', 'sentence'],
num_rows: 1700
})
test: Dataset({
features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/bemba-speech-csikasote.bemba_train_dev_sets_processedbemba-french_sentence-pairs
Bemba-French_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-French_Sentence-Pairs
File Size: 79644493 bytes
Languages: Bemba, French
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-french_sentence-pairs.Code-170k-bemba
Dataset Description
Code-170k-bemba is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Bemba, making coding education accessible to Bemba speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Bemba language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-bemba.bembaspeech_plus_jw_processed
