datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
my-fonts-databasefon-code-switching-evaluation
French-Fon Code-Switching Evaluation Benchmark
Overview
This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios.
The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon.
The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.Fon_French_Daily_Dialogues_Parallel_Data
[!NOTE]
Dataset origin: https://zenodo.org/records/4432712
Description
We aim to collect, clean, and store corpora of Fon and French sentences for Natural Language Processing researches including Neural Machine Translation, Named Entity Recognition, etc. for Fon, a very low-resourced and endangered African native language.
Fon (also called Fongbe) is an African-indigenous language spoken mostly in Benin, Togo, and Nigeria - by about 2 million people.
As training data is crucial to… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Fon_French_Daily_Dialogues_Parallel_Data.fon-swati_sentence-pairs
Fon-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Swati_Sentence-Pairs
Number of Rows: 31019
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-swati_sentence-pairs.fon-lingala_sentence-pairs
Fon-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Lingala_Sentence-Pairs
Number of Rows: 69800
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-lingala_sentence-pairs.fon-hausa_sentence-pairs
Fon-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Hausa_Sentence-Pairs
Number of Rows: 103601
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-hausa_sentence-pairs.bemba-fon_sentence-pairs
Bemba-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Fon_Sentence-Pairs
Number of Rows: 86130
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-fon_sentence-pairs.bambara-fon_sentence-pairs
Bambara-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Fon_Sentence-Pairs
Number of Rows: 25525
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-fon_sentence-pairs.fon-tumbuka_sentence-pairs
Fon-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tumbuka_Sentence-Pairs
Number of Rows: 73794
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tumbuka_sentence-pairs.fon-wolof_sentence-pairs
Fon-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Wolof_Sentence-Pairs
Number of Rows: 23879
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-wolof_sentence-pairs.fon-umbundu_sentence-pairs
Fon-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Umbundu_Sentence-Pairs
Number of Rows: 56632
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-umbundu_sentence-pairs.fon-twi_sentence-pairs
Fon-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Twi_Sentence-Pairs
Number of Rows: 87200
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-twi_sentence-pairs.fon-kamba_sentence-pairs
Fon-Kamba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Kamba_Sentence-Pairs
Number of Rows: 35955
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-kamba_sentence-pairs.fon-kikuyu_sentence-pairs
Fon-Kikuyu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Kikuyu_Sentence-Pairs
Number of Rows: 34694
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-kikuyu_sentence-pairs.fon-tsonga_sentence-pairs
Fon-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tsonga_Sentence-Pairs
Number of Rows: 97374
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tsonga_sentence-pairs.fon-kimbundu_sentence-pairs
Fon-Kimbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Kimbundu_Sentence-Pairs
Number of Rows: 44269
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-kimbundu_sentence-pairs.fon-fulah_sentence-pairs
Fon-Fulah_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Fulah_Sentence-Pairs
Number of Rows: 81491
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-fulah_sentence-pairs.fon-xhosa_sentence-pairs
Fon-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Xhosa_Sentence-Pairs
Number of Rows: 90621
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-xhosa_sentence-pairs.fon-tigrinya_sentence-pairs
Fon-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tigrinya_Sentence-Pairs
Number of Rows: 59562
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tigrinya_sentence-pairs.fon-pedi_sentence-pairs
Fon-Pedi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Pedi_Sentence-Pairs
Number of Rows: 58696
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-pedi_sentence-pairs.dyula-fon_sentence-pairs
Dyula-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Fon_Sentence-Pairs
Number of Rows: 41426
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-fon_sentence-pairs.fon-oromo_sentence-pairs
Fon-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Oromo_Sentence-Pairs
Number of Rows: 50665
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-oromo_sentence-pairs.fon-nuer_sentence-pairs
Fon-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Nuer_Sentence-Pairs
Number of Rows: 12128
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-nuer_sentence-pairs.fon-igbo_sentence-pairs
Fon-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Igbo_Sentence-Pairs
Number of Rows: 55155
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-igbo_sentence-pairs.Legacy-Font-and-Romanized-Tamil-Corpusfon-zulu_sentence-pairs
Fon-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Zulu_Sentence-Pairs
Number of Rows: 137778
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-zulu_sentence-pairs.fon-rundi_sentence-pairs
Fon-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Rundi_Sentence-Pairs
Number of Rows: 89002
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-rundi_sentence-pairs.fon-kinyarwanda_sentence-pairs
Fon-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Kinyarwanda_Sentence-Pairs
Number of Rows: 105316
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-kinyarwanda_sentence-pairs.fon-kongo_sentence-pairs
Fon-Kongo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Kongo_Sentence-Pairs
Number of Rows: 57877
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-kongo_sentence-pairs.fon-chichewa_sentence-pairs
Fon-Chichewa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Chichewa_Sentence-Pairs
Number of Rows: 131151
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-chichewa_sentence-pairs.
