Team Ai
Datasetpublic

RayNene/adaption-agronomy-qa-pairs

East Africa Agronomy QA Agricultural AI for the languages and people feeding East Africa East Africa's food system is enormous. Agriculture remains at the centre of the region's economy and livelihoods, while Africa as a whole has around 300 million people employed in agrifood systems and the highest regional share of employment in agrifood systems globally, at 64.5%. Agriculture accounts for roughly 74.4% of Africa's agrifood-system employment.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/RayNene/adaption-agronomy-qa-pairs.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes47downloads
Dataset Card

East Africa Agronomy QA

Agricultural AI for the languages and people feeding East Africa

East Africa's food system is enormous. Agriculture remains at the centre of the region's economy and livelihoods, while Africa as a whole has around 300 million people employed in agrifood systems and the highest regional share of employment in agrifood systems globally, at 64.5%. Agriculture accounts for roughly 74.4% of Africa's agrifood-system employment. :contentReference[oaicite:1]{index=1}

At the same time, the next generation is growing up in a region where farming has to become more productive, climate-resilient, and technology-enabled. That creates a simple problem for AI: agricultural knowledge is valuable, but much of the world's agricultural AI is developed around English and high-resource languages.

This project is built to push in the opposite direction.

The dataset brings agricultural knowledge into six languages used across East and Central Africa: Swahili, Kinyarwanda, Somali, Lingala, Luganda, and English, opening agricultural AI research to a potential audience of 100M+ farmers and agricultural communities across the language regions represented.

The goal is simple: agricultural AI should be able to speak to farmers in the languages they actually use.

banner

Dataset

This is a remastered agricultural instruction-tuning dataset containing 14,996 question-and-answer pairs focused on practical agronomy and agricultural science.

The dataset covers areas including:

  • โ€”๐ŸŒฑ Soil management
  • โ€”๐ŸŒพ Crop cultivation
  • โ€”๐Ÿ› Pest control
  • โ€”๐Ÿ’ง Water and growing systems
  • โ€”๐ŸŒฟ Pasture improvement
  • โ€”๐Ÿฅฌ Hydroponics
  • โ€”๐ŸŒฑ Cassava farming
  • โ€”๐Ÿ„ Livestock and agricultural systems

The objective is not to create another collection of textbook definitions. It is to build training data around the kinds of agricultural questions that can arise in real farming and agricultural decision-making.

Language Coverage

The project is designed around:

LanguageRole
SwahiliMajor East African lingua franca
KinyarwandaRwanda and surrounding Great Lakes communities
SomaliSomali-speaking agricultural communities
LingalaCentral African agricultural communities
LugandaUganda and surrounding communities
EnglishHigh-resource reference and scientific language

This multilingual direction matters because agricultural knowledge often travels through extension officers, community organisations, research institutions, and informal networks rather than directly from academic literature to farmers.

A model that can reason about agriculture in English but cannot communicate the same knowledge naturally in Swahili, Kinyarwanda, Somali, Lingala, or Luganda leaves a significant part of the problem unsolved.

Why Agriculture?

Agrifood systems are not a peripheral part of Africa's economy.

FAO estimates that around 300 million people in Africa work in agrifood systems, with agriculture accounting for the majority of those jobs. :contentReference[oaicite:2]{index=2}

And the importance extends beyond people formally classified as farmers. Agrifood systems connect production to livestock, inputs, processing, transportation, trade, and food services.

The opportunity is therefore much larger than "AI for farmers."

It is AI for the agricultural knowledge layer of an entire food system.

What the Model Should Eventually Do

A useful agricultural model should be able to handle questions such as:

  • โ€”What might be causing poor crop growth?
  • โ€”How can soil fertility be improved?
  • โ€”What are practical approaches to pest management?
  • โ€”How should a farmer think about water management?
  • โ€”What cultivation practices are appropriate for a particular crop?
  • โ€”How can agricultural practices adapt to changing environmental conditions?

And importantly, it should be able to answer these questions in the user's language.

Agricultural AI Evaluation

Agricultural LLM evaluation is still an emerging field, which makes domain-specific benchmarks particularly useful.

Relevant evaluation work includes AgriBench, which evaluates agricultural capabilities in multimodal models across real agricultural topics, and AgriEval, which evaluates agricultural knowledge across six major categories and 29 subcategories, including memorisation, understanding, inference, and generation. :contentReference[oaicite:3]{index=3}

The dataset is intended to support evaluation against these kinds of agricultural reasoning tasks, alongside multilingual and domain-specific evaluations.

Benchmark results should be reported only where the model has actually been evaluated on the corresponding benchmark.

Dataset Statistics

14,996 instruction-tuning examples

Domain

DomainShare
Agriculture96%
Science2%
Animal & Nature2%

Response style

ToneShare
Practical58%
Informative32%
Explanatory6%

The emphasis on practical responses is deliberate: agricultural AI should help communicate actionable knowledge rather than simply reproduce academic descriptions.

Remastering Results

The dataset was remastered using Adaption's Adaptive Data platform.

The remastered dataset achieved a B quality grade, representing a reported 46.0% relative quality improvement.

Quality Gains

<img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/ef5ec192-6220-4a8b-8f60-82bf042fc46b.png" alt="QualityGains" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />

Grade Improvement

<img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/3fe64049-6641-43de-be09-0923f4597419.png" alt="Grade" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />

Percentile Chart

<img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/bc847531-d403-4095-94d7-41b2d89fae36.png" alt="Percentile Chart" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />

Intended Use

This dataset is intended for:

  • โ€”๐ŸŒพ Agricultural LLM fine-tuning
  • โ€”๐Ÿค– Agricultural question-answering systems
  • โ€”๐Ÿง‘โ€๐ŸŒพ Farmer advisory assistants
  • โ€”๐Ÿ“š Agricultural education
  • โ€”๐Ÿ”ฌ Agronomy and NLP research
  • โ€”๐ŸŒ African-language AI research
  • โ€”๐ŸŒฆ๏ธ Climate-smart agriculture applications
  • โ€”๐Ÿ“Š Agricultural model evaluation

The Bigger Picture

East Africa does not need AI that simply knows what a crop is.

It needs AI that can help make agricultural knowledge accessible, local, multilingual, and useful.

The long-term direction is a generation of agricultural models that can combine agronomy with local context and communicate through the languages already used by farming communities.

From agricultural research to agricultural knowledge. From agricultural knowledge to agricultural action.

This dataset is a step toward that infrastructure.

Responsible Use

Agricultural recommendations can directly affect livelihoods. Models trained on this dataset should therefore be treated as decision-support tools rather than replacements for agricultural extension officers, agronomists, veterinarians, or local expertise.

Recommendations should account for local climate, soil conditions, crop varieties, pests, regulations, and farming practices.

Acknowledgements

This dataset was remastered using Adaption's Adaptive Data platform as part of the Adaption Labs AutoScientist Challenge.

Citation

bibtex
@dataset{east_africa_agronomy_2026,
  title     = {East Africa Agronomy QA Dataset},
  author    = {Nene, Ray},
  year      = {2026},
  publisher = {Hugging Face}
}