Team Ai
Apppublic

Zoe/GS-QA-Leaderboard

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
App README

GS-QA Leaderboard

This static Space presents results for GS-QA2, a benchmark containing 3,300 geospatial question-answer pairs across 53 templates: 28 vector-only templates and 25 raster or raster-vector templates.

Citation

Zhuocheng Shang, Shahd Elmahallawy, Zabir Al Nazi, Vagelis Hristidis, and Ahmed Eldawy. 2026. GS-QA2: A Benchmark for Question Answering over Raster–Vector Data. In Proceedings of the 34th ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26). Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3841645.3843431

bibtex
@inproceedings{shang2026gsqa2,
  author    = {Zhuocheng Shang and Shahd Elmahallawy and Zabir Al Nazi and
               Vagelis Hristidis and Ahmed Eldawy},
  title     = {{GS-QA2}: A Benchmark for Question Answering over Raster--Vector Data},
  booktitle = {Proceedings of the 34th ACM International Conference on Advances
               in Geographic Information Systems},
  series    = {SIGSPATIAL '26},
  year      = {2026},
  month     = nov,
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  doi       = {10.1145/3841645.3843431},
  isbn      = {979-8-4007-2950-8}
}

Metrics

GS-QA uses output-specific evaluation:

  • —Vector entity names: parsed token F1 on attempted answers (higher is better).
  • —Vector locations: strict accuracy with geodesic error at most 5 meters.
  • —Vector directions: strict accuracy with circular angular error at most 5 degrees.
  • —Vector distance, length, area, and count: strict accuracy with relative error at most 5%.
  • —Raster and raster–vector templates: strict per-question accuracy using the output-specific tolerances from Table 10. Elevation error must be at most 10 meters; elevation-coverage error at most 5 percentage points; slope or aspect error at most 5 degrees; ruggedness and other numeric relative error at most 5%; entity-name token F1 at least 0.8; and threshold labels must match exactly. Every component of a compound answer must pass.

Template IDs identify the benchmark track: V1–V28 are vector-only, R1–R11 are raster-only, and VR1–VR14 are raster–vector. Raster-related results are also grouped by local, focal, zonal, and global map-algebra operations.

Entity quality and capped numeric relative error are averaged only over attempted questions. Strict accuracy uses every benchmark question, including failures, matching Tables 11-14 of the GS-QA2 experiment paper.

Raster operation labels follow the template taxonomy in Tables 5-7. In particular, R10, R11, VR5, and VR6 are Local operations. Table 14 omits the VR6 detail row, but its reported overall averages include all 25 raster-related templates and imply zero accuracy for VR6 across the three paper systems.

The leaderboard deliberately avoids a cross-family composite score because the answer families use different units and have different error semantics.

Submitting a result

There are two participation paths:

  1. 1.Self-service vector: clone this Space, follow runner/docker/README.md, and run the 2,800-question vector track locally or in the submitter's own Hugging Face Job. The lightweight runner image is approximately 190 MB and does not require PostGIS.
  2. 2.Managed raster/full: open a Community discussion with the model, method, exact revisions, generation settings, and inference requirements. Maintainers run the 500-question raster track or complete 3,300-question benchmark against the reference PostGIS/PostGIS-Raster database on the lab cluster.

For the self-service path, publish the complete predictions.jsonl in a public Hugging Face dataset repository. Maintainers score it with the reference evaluator and create the final submission.json plus per-question evaluation CSVs. Self-reported summary scores are not accepted directly.

Maintainers run submission/validate_submission.py before importing a run. Self-reported summary rows are not accepted directly into the official ranking. See submission/README.md for the complete workflow.

Reproducing the data

The live leaderboard is generated from Tables 11-14 of the GS-QA2 experiment paper:

bash
cd huggingface/leaderboard
python build_paper_results.py

The generated data/results.json powers the public static app. Both JSON and CSV are downloadable from the leaderboard and bundled with the Space so the published results remain tied to a specific Space revision. The 3,300 benchmark questions remain in the versioned Zoe/GS-QA2 dataset; the Space links to that dataset instead of duplicating its larger artifacts.

For future experiment artifacts, build_table6_metrics.py and prepare_results.py apply the same attempted-question, capped-error, and strict accuracy rules before a verified run is merged.