michael7ma/ogd4all-benchmark
OGD4All Benchmark This is a 199-question benchmark that was used to evaluate the overall performance of OGD4All and different configurations (LLM, orchestration, ...). OGD4All is an LLM-based prototype system enabling an easy-to-use, transparent interaction with Geospatial Open Government Data through natural language. Each question requires GIS, SQL and/or topological operations on zero, one, or multiple datasets in GPKG or CSV formats to be answered. Tasks The… See the full description on the dataset page: https://huggingface.co/datasets/michael7ma/ogd4all-benchmark.
OGD4All Benchmark 
This is a 199-question benchmark that was used to evaluate the overall performance of OGD4All and different configurations (LLM, orchestration, ...). OGD4All is an LLM-based prototype system enabling an easy-to-use, transparent interaction with Geospatial Open Government Data through natural language. Each question requires GIS, SQL and/or topological operations on zero, one, or multiple datasets in GPKG or CSV formats to be answered.
Tasks
The benchmark can be used to evaluate systems against two main tasks:
- Dataset Retrieval: Given actual metadata of 430 City of Zurich datasets and a question, identify the subset of $k$ relevant datasets to answer the question. Note that the case $k=0$ is included
- Dataset Analysis: Given a set of relevant datasets, corresponding metadata and a question, appropriately process these datasets (e.g. via generated Python code snippets) and produce a textual answer. Note that OGD4All can accompany this answer with an interactive map, plots and/or tables, but only the textual answer is evaluated.
Evaluation
Metrics
Dataset Retrieval
To evaluate dataset retrieval, rely on the "relevant_datasets" list in the "outputs" dict, which gives you the list of relevant titles. You can map between metadata files and titles using the data/dataset_title_to_file.csv.
Dataset Analysis
To evaluate dataset analysis, provide your architecture with relevant datasets specified in the "outputs" dict and the question, then either manually compare the generated answer with the ground-truth answer in the "outputs" dict, or use the LLM judge system prompt given in eval_prompts/LLM_JUDGE_SYSTEM_PROMPT.txt, with the question, reference and predicted answer provided via a subsequent user message.
[!NOTE] A few questions were found to have multiple options of valid relevant datasets, and also multiple valid answers. Therefore, your evaluation should consider the attributesalternative_relevant_datasetsandalternative_answerif present.
Benchmark Notes
- benchmark_german.jsonl is the main benchmark, developed in German. All metadata/datasets are always in German.
- We further provide automatically-translated versions of the questions (via DeepL API) in benchmarkenglish.jsonl, benchmarkfrench.jsonl, and benchmark_italian.jsonl.
- benchmark_template.jsonl is the template that was used for generating the previously mentioned benchmark, with templated questions that can be instantiated with different arguments.
benchmarks/gt_scriptscontains Python files that were hand-written to generate the ground-truth answer for each question that has relevant datasets. The filename corresponds to the benchmark entry ID.- The city of Zurich datasets are under CC-0 license. Recent versions can be downloaded here, but for evaluation you should use the included datasets, as some answers might change otherwise. NOTE: if you only wish to evaluate the "Dataset Analysis" task, you can download only the subset of datasets referenced as "relevant_datasets" in the benchmark entries.
- For "Dataset Analysis", it makes sense to equip the agent/LLM with a geocoding utility. OGD4All relies on Google's Geocoding API.
Citation
If you use this benchmark in your research, gladly cite our accompanying paper:
@article{siebenmann_ogd4all_2025,
archivePrefix = {arXiv},
arxivId = {2602.00012},
author = {Siebenmann, Michael and S{\'a}nchez-Vaquerizo, Javier Argota and Arisona, Stefan and Samp, Krystian and Gisler, Luis and Helbing, Dirk},
journal = {arXiv preprint arXiv:2602.00012},
month = {nov},
title = {{OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Models}},
url = {https://arxiv.org/abs/2602.00012},
year = {2025}
}Next to the benchmark, this paper (accepted at IEEE CAI 2026) introduces the OGD4All architecture, which achieves high recall and correctness scores, even with "older" frontier models such as GPT-4.1. OGD4All's source code is publicly available: https://github.com/ethz-coss/ogd4all
