Atnafu/Afri-MCQA
Afri-MCQA: Multimodal Cultural Question Answering for African Languages Paper Overview Afri-MCQA is the first multilingual cultural question-answering benchmark covering 8k Q&A pairs across 16 African languages from 13 countries. The benchmark offers parallel English-African language Q&A pairs across text and speech modalities, entirely created by native speakers. Supported Tasks Visual Question Answering (VQA): Multiple-choice and open-ended QA… See the full description on the dataset page: https://huggingface.co/datasets/Atnafu/Afri-MCQA.
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
Overview
Afri-MCQA is the first multilingual cultural question-answering benchmark covering 8k Q&A pairs across 16 African languages from 13 countries. The benchmark offers parallel English-African language Q&A pairs across text and speech modalities, entirely created by native speakers.
Supported Tasks
- Visual Question Answering (VQA): Multiple-choice and open-ended QA grounded in culturally relevant images
- Visual Audio Question Answering: Speech-based QA grounded in culturally relevant images in native African languages and African-accented English
- Language Identification (LID): Identifying which of the 15 languages is spoken
- Automatic Speech Recognition (ASR): Transcribing spoken African language audio
Languages
Total speaker population: ~412.6 million
Dataset Structure
Each sample includes:
- image: Culturally relevant image
- question_english / question_native: Question in both languages
- options_english / options_native: Four multiple-choice options
- answer: Correct answer
- audio_question_english / audio_question_native: Audio recordings
- category: One of 10 cultural categories
- country / language: Origin metadata
Cultural Categories
🏛️ Geography & Landmarks | 👤 Public Figures & Pop Culture | 🍲 Cooking & Food | 👕 Objects & Clothing |🎭 Traditions & History | 🏢 Brands & Companies | 🌿 Plants & Animals | 👨👩👧 People & Everyday Life | 🚗 Vehicles & Transportation | ⚽ Sports & Recreation
Licensing
This dataset is released under CC-BY-NC-4.0.
Usage
from datasets import load_dataset
# Per language
amharic_dev = load_dataset("Atnafu/Afri-MCQA", "Amharic_dev", split="dev")
amharic_test = load_dataset("Atnafu/Afri-MCQA", "Amharic_test", split="test")
# All languages grouped
all_dev = load_dataset("Atnafu/Afri-MCQA", "all_dev", split="dev")
all_test = load_dataset("Atnafu/Afri-MCQA", "all_test", split="test")Citation
@inproceedings{tonja-etal-2026-afri,
title = "Afri-{MCQA}: Multimodal Cultural Question Answering for {A}frican Languages",
author = "Tonja, Atnafu Lambebo and
Anand, Srija and
Villa-Cueva, Emilio and
Azime, Israel Abebe and
Alabi, Jesujoba Oluwadara and
Mohamed, Muhidin A. and
Yadeta, Debela Desalegn and
Abadi, Negasi Haile and
Oppong, Abigail and
Obiefuna, Nnaemeka Casmir and
Abdulmumin, Idris and
Etori, Naome A and
Wairagala, Eric Peter and
Tshinu, Kanda Patrick and
Emmanuel, Imanigirimbabazi and
Malema, Gabofetswe and
Aji, Alham Fikri and
Adelani, David Ifeoluwa and
Solorio, Thamar",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.1869/",
doi = "10.18653/v1/2026.acl-long.1869",
pages = "40249--40282",
ISBN = "979-8-89176-390-6",
abstract = "Africa is home to over one-third of the world{'}s languages, yet remains severely underrepresented in multimodal AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering benchmark containing 7.5k Q A pairs across 15 African languages from 12 countries. The benchmark offers parallel text and speech modalities and was entirely created by native speakers. We find that models show poor performance across evaluated cultures, with near-zero accuracy on open-ended VQA when queried through native language or speech. To test linguistic competence, we include control experiments meant to assess this specific aspect separate from cultural knowledge, and we observe significant performance gaps between native languages and English for both text and speech. These findings underscore the pressing need for speech-first approaches, culturally grounded pretraining, and cross-lingual cultural transfer. We release Afri-MCQA to support more inclusive multimodal AI development."
}