SyedNazmusSakib/PlantInquiryVQA
PlantInquiryVQA — Thinking Like a Botanist Benchmark and framework for multi-turn, intent-driven visual question answering in plant pathology. Accepted at ACL 2026 Findings. Overview PlantInquiryVQA formalises diagnostic reasoning in plant pathology as a Chain-of-Inquiry (CoI) — an ordered sequence of visually-grounded questions that adapts to the plant's severity and the expert's epistemic intent (Diagnosis / Prognosis / Management). The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/SyedNazmusSakib/PlantInquiryVQA.
PlantInquiryVQA — Thinking Like a Botanist
Benchmark and framework for multi-turn, intent-driven visual question answering in plant pathology.
Accepted at ACL 2026 Findings.
    
Overview
PlantInquiryVQA formalises diagnostic reasoning in plant pathology as a Chain-of-Inquiry (CoI) — an ordered sequence of visually-grounded questions that adapts to the plant's severity and the expert's epistemic intent (Diagnosis / Prognosis / Management).
The benchmark evaluates whether modern Multimodal Large Language Models (MLLMs) can reason like a botanist, not just classify a leaf. Key findings from benchmarking 18 state-of-the-art models:
- All 18 MLLMs describe symptoms competently but fail at reliable clinical reasoning (top Clinical Utility score = 0.188 / 1.0)
- Structured question-guided inquiry improves diagnostic accuracy by ~48% over direct diagnosis
- Structured CoI reduces hallucination significantly compared to free-form dialogue
Dataset at a Glance
Covered crop species (34 total)
Apple · Arabian Jasmine · Bitter Gourd · Blueberry · Bottle Gourd · Cauliflower · Cherry · Corn · Cotton · Cucumber · Eggplant/Brinjal · Grape · Guava · Hibiscus · Jackfruit · Lemon · Litchi · Mango · Orange · Papaya · Peach · Peas · Pepper · Pepper Bell · Potato · Raspberry · Rice · Rubber · Soybean · Squash · Strawberry · Sunflower · Tea · Tomato
Quick Load
from datasets import load_dataset
# Load train / test splits (metadata only — no images)
ds = load_dataset("SyedNazmusSakib/PlantInquiryVQA")
train = ds["train"]
test = ds["test"]
print(train[0])
# {
# 'image_id': 'f650d82227e534b8.jpg',
# 'crop': 'Bottle Gourd',
# 'disease': 'healthy',
# 'category': 'healthy',
# 'severity': '',
# 'question': 'What crop is shown in this image?',
# 'answer': 'This leaf is from a Bottle Gourd plant ...',
# 'question_category': 'crop_identification',
# 'visual_grounding': '',
# 'question_number': 1.0,
# 'dataset_source': 'non_disease'
# }Load with images
Images live in the images/ folder of this repository, named by image_id.
from datasets import load_dataset
from huggingface_hub import hf_hub_download
from PIL import Image
ds = load_dataset("SyedNazmusSakib/PlantInquiryVQA", split="test")
def attach_image(row):
img_path = hf_hub_download(
repo_id="SyedNazmusSakib/PlantInquiryVQA",
repo_type="dataset",
filename=f"images/{row['image_id']}",
)
row["image"] = Image.open(img_path).convert("RGB")
return row
# Attach images on demand (lazy)
sample = attach_image(ds[0])Download the full image corpus locally
# Install helper
pip install huggingface_hub
# Download everything to ./images/
python -c "
from huggingface_hub import snapshot_download
snapshot_download(
repo_id='SyedNazmusSakib/PlantInquiryVQA',
repo_type='dataset',
local_dir='./PlantInquiryVQA',
allow_patterns=['images/*', 'data/*.csv', 'visual_cues/*', 'diseases_knowledge_base/*'],
)
"Or use the provided script (after cloning the GitHub repo):
python scripts/download_images.pyRepository Structure
SyedNazmusSakib/PlantInquiryVQA (HuggingFace)
├── README.md ← this file (dataset card)
├── CITATION.cff ← machine-readable citation
├── requirements.txt
│
├── data/
│ ├── train.csv ← 82,800 QA rows (80% split)
│ └── test.csv ← 55,268 QA rows (20% split)
│
├── images/ ← 24,950 leaf JPEGs (~3.5 GB)
│ ├── 00009faac7cf68de.jpg
│ └── ...
│
├── diseases_knowledge_base/
│ ├── all_cards.jsonl ← 116-disease expert knowledge cards
│ └── <crop>/<disease>.json
│
└── visual_cues/
└── visual_cues.json ← 24,950 × expert-verified visual cuesCSV Schema
Evaluation Protocols
Benchmark Results (18 MLLMs)
Best-in-class per metric (full table in paper Table 2):
Models benchmarked include: Gemini 3 Flash/Pro, Claude (via OpenRouter), GPT-4o, Qwen3-VL (8B/32B/235B), Qwen2.5-VL (7B/32B/72B), LLaMA-3.2 (11B/90B), LLaMA-4 Maverick, Grok-4.1-Fast, Pixtral-12B, Mistral Medium 3.1, Mistral Small 24B, Ministral (3B/8B), Nemotron-12B, Phi-4-Multimodal, Seed-1.6-Flash.
Domain-Specific Metrics
Defined in paper Appendix A.1:
- S_dis — Disease Identification Score (strict entity match)
- S_safe — Safety Score (false-reassurance penalty)
- S_clin = 0.5·Sdis + 0.3·Sact − 0.2·(1 − S_safe) — composite clinical utility
- S_vg — Visual Grounding recall of expert-verified cues
- E — Explainability Efficiency (verified cues per 100 words)
- B — Prevalence Bias (Eq. 7)
- F — Cross-Class Fairness (Eq. 8)
Supplementary Files
diseases_knowledge_base/
Expert disease cards for each of the 116 disease categories across 34 crops. Each card contains:
- Disease description and causal agent
- Visual diagnostic criteria
- Severity progression markers
- Management recommendations
visual_cues/visual_cues.json
24,950-entry lookup table mapping each image_id to expert-verified visual cues used for visual grounding evaluation.
Reproducing Results
git clone https://github.com/SyedNazmusSakib/PlantInquiryVQA
cd PlantInquiryVQA
pip install -r requirements.txt
cp .env.example .env # fill in API keys
# Download images from this HF repo
python scripts/download_images.py
# Run evaluation (Guided setting)
python eval/test_1_gemini3_flash.py
# Aggregate all results
python eval/compute_cascading_all_models.py
python eval/compute_fairness_all_models.pyLicence
Citation
If you use PlantInquiryVQA, please cite:
@article{sakib2026thinking,
title={Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry},
author={Sakib, Syed Nazmus and Haque, Nafiul and Amin, Shahrear Bin and Abdullah, Hasan Muhammad and Hasan, Md Mehedi and Hossain, Mohammad Zabed and Arman, Shifat E},
journal={arXiv preprint arXiv:2604.20983},
year={2026}
}Contact: Open an issue on GitHub or reach out to the corresponding author listed in the paper.
We thank Ali Akbar for large-scale data collection and Abdullah Shahriar for creating the figures and diagrams.
