fai-adh/fon-code-switching-evaluation
French-Fon Code-Switching Evaluation Benchmark Overview This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios. The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon. The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.
French-Fon Code-Switching Evaluation Benchmark
Overview
This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios.
The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon.
The benchmark was developed as part of an academic research project on the evaluation of language models in multilingual and low-resource language contexts.
Research Question
The main research question is:
How well can a small open-weight language model maintain contextual understanding when users switch between French and Fon within the same interaction?
Model Evaluated
The benchmark was evaluated using:
Qwen/Qwen3-4B-Instruct-2507
The model contains approximately 4 billion parameters and was selected because it is an open-weight language model within the required range for this study.
Benchmark Design
The benchmark contains:
- 40 base scenarios
- 4 experimental conditions
- 160 total evaluation instances
The 40 scenarios are distributed across four categories:
- Daily-life situations
- Reasoning
- Context tracking
- Local and cultural context
Each category contains 10 base scenarios.
Experimental Conditions
Each scenario was evaluated under four conditions.
1. French baseline (fr)
The complete task is presented in French.
This condition provides a baseline for measuring the model's performance in French.
2. Fon-only (fon)
The relevant information and task are presented in Fon.
This condition evaluates the model's ability to understand and use information expressed in Fon.
3. French-Fon code-switching (fr_fon)
French and Fon are combined within the same interaction.
This condition evaluates the model's ability to maintain contextual understanding when switching between the two languages.
4. Critical information in Fon (critical_fon)
The final question is presented in French, while the information necessary to answer the question is provided only in Fon.
This condition specifically tests whether the model can identify, understand and reason over critical information expressed in Fon.
Evaluation
Each model response was manually evaluated using a binary correctness criterion:
1= correct0= incorrect
The evaluation primarily considered whether the model produced the expected semantic answer.
A qualitative error taxonomy was also used to analyze incorrect responses.
Error Taxonomy
The following error categories were defined:
- E1 — Wrong Fon understanding: incorrect interpretation of information expressed in Fon.
- E2 — Information loss: relevant contextual information is lost or modified.
- E3 — Code-switching interpretation error: incorrect interpretation of the interaction when French and Fon are combined.
- E4 — Reasoning error: the relevant information is available but the model fails to perform the required reasoning.
- E5 — Hallucination: the model introduces unsupported information.
- E6 — Partially correct: the response contains part of the expected information but is incomplete.
- E7 — Refusal or incomprehension: the model states that it cannot understand or answer the task.
- E8 — Other: errors that do not fit the previous categories.
Translation Verification
The 40 Fon translations used in the benchmark were manually verified before the final evaluation.
The translations were considered valid and no correction was required.
Automatic translation was therefore not treated as the gold standard for evaluation.
Results
The benchmark produced the following overall accuracy:
31.9%
Performance by experimental condition:
Performance by scenario category:
Error Analysis
Among the 109 incorrect responses, the main observed error types were:
The two dominant error types were E1 and E3, which together represented approximately 70.6% of all observed errors.
The results show a substantial degradation when Fon is involved in the task compared with the French baseline.
The French baseline reached 97.5% accuracy, whereas the Fon-only and French-Fon conditions reached 7.5% and 5.0%, respectively.
The critical-information-in-Fon condition reached 17.5%.
These results support the hypothesis that the model experiences substantial difficulty maintaining contextual understanding when information is expressed in Fon.
However, the hypothesis that the critical-information-in-Fon condition would necessarily produce the lowest performance was not confirmed, since the French-Fon condition obtained the lowest accuracy.
Limitations
This benchmark has several limitations.
First, it contains 160 evaluation instances and therefore represents a relatively small evaluation set.
Second, the benchmark focuses on French and Fon and should not be generalized automatically to all African languages or multilingual settings.
Third, the evaluation uses manually constructed scenarios and therefore reflects the design choices of the researchers.
Finally, the results concern the evaluated model and benchmark conditions and should not be interpreted as a general statement about all language models or about the overall ability of AI systems to understand Fon.
Intended Use
This dataset can be used for:
- Evaluating multilingual language models.
- Studying code-switching between French and Fon.
- Studying low-resource language understanding.
- Analyzing contextual understanding in multilingual interactions.
- Reproducing the experiments described in the associated research project.
Dataset Structure
The repository contains evaluation scenarios, model responses, annotations, result tables, visualizations and experimental notebooks.
Citation
If you use this benchmark in academic work, please cite the associated research project.
License
This dataset is released under the MIT License.
