datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
MAP-CC
MAP-CC
🌐 Homepage | 🤗 MAP-CC | 🤗 CHC-Bench | 🤗 CT-LLM | 📖 arXiv | GitHub
An open-source Chinese pretraining dataset with a scale of 800 billion tokens, offering the NLP community high-quality Chinese pretraining data.
Disclaimer
This model, developed for academic purposes, employs rigorously compliance-checked training data to uphold the highest standards of integrity and compliance. Despite our efforts, the inherent complexities of data and the broad spectrum of… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MAP-CC.war3-maps
Warcraft III Community Map Archive
This public dataset preserves community-created Warcraft III maps and campaigns
for interoperability testing, search, research, and long-term access. Files are
deduplicated by SHA-256. Titles and other metadata are extracted with
war3-manager where the format permits.
Search and download individual maps: https://war3-archive.github.io/war3-maps/
Source and issue tracker: https://github.com/war3-archive/war3-maps
Layout… See the full description on the dataset page: https://huggingface.co/datasets/magicwenli/war3-maps.map-spatial-benchmark
Map-based Spatial Reasoning Benchmark
A multi-view map-based spatial reasoning benchmark. Each row is one multiple-choice
question instance over a registered map image; models must answer with a single option
letter. Four tasks (T1–T4), four base-map views, and controlled evidence conditions
(direct / query / oracle) and world perturbations (transform / world layers) allow
fine-grained analysis of spatial reasoning robustness.
Task overview
Task
Question… See the full description on the dataset page: https://huggingface.co/datasets/mapspatial/map-spatial-benchmark.MAPBench-V1For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
MapEval-Visual
MapEval-Visual
This dataset was introduced in MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
Example
Query
I am presently visiting Mount Royal Park . Could you please inform me about the nearby historical landmark?
Options
Circle Stone
Secret pool
Maison William Caldwell Cottingham
Poste de cavalerie du Service de police de la Ville de Montreal
Correct Option
Circle Stone
Prerequisite
Download… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-Visual.MapRoute-BaselineIV-Bench
IV‑Bench: A Benchmark for Image‑Grounded Video Perception and Reasoning in Multimodal LLMs
Dataset Availability
Due to privacy policy, only a subset of the IV‑Bench dataset is publicly available. Specifically, we release 1,680 samples, including video IDs, image–text queries, and distractors.
Usage
Detailed usage instructions can be found on GitHub: IV‑Bench
