Team Ai
Datasetpublic

TankVincent/MapEval-Textual

MapEval-Textual MapEval-Textual is created using MapQaTor. Usage from datasets import load_dataset # Load dataset ds = load_dataset("MapEval/MapEval-Textual", name="benchmark") # Generate better prompts for item in ds["test"]: # Start with a clear task description prompt = ( "You are a highly intelligent assistant. " "Based on the given context, answer the multiple-choice question by selecting the correct option.\n\n" "Context:\n"… See the full description on the dataset page: https://huggingface.co/datasets/TankVincent/MapEval-Textual.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes9downloads
Dataset Card

MapEval-Textual

MapEval-Textual is created using MapQaTor.

Usage

python
from datasets import load_dataset

# Load dataset
ds = load_dataset("MapEval/MapEval-Textual", name="benchmark")

# Generate better prompts
for item in ds["test"]:
    # Start with a clear task description
    prompt = (
        "You are a highly intelligent assistant. "
        "Based on the given context, answer the multiple-choice question by selecting the correct option.\n\n"
        "Context:\n" + item["context"] + "\n\n"
        "Question:\n" + item["question"] + "\n\n"
        "Options:\n"
    )
    
    # List the options more clearly
    for i, option in enumerate(item["options"], start=1):
        prompt += f"{i}. {option}\n"
    
    # Add a concluding sentence to encourage selection of the answer
    prompt += "\nSelect the best option by choosing its number."

    # Use the prompt as needed
    print(prompt)  # Replace with your processing logic

Leaderboard

ModelOverallPlace InfoNearbyRoutingTripUnanswerable
Claude-3.5-Sonnet66.3373.4473.4975.7649.2540.00
Gemini-1.5-Pro66.3365.6374.7069.7047.7685.00
GPT-4o63.3364.0674.7069.7049.2540.00
GPT-4-Turbo62.3367.1971.0871.2147.7630.00
Gemini-1.5-Flash58.6762.5067.4766.6738.8150.00
GPT-4o-mini51.0046.8863.8657.5840.3025.00
GPT-3.5-Turbo37.6726.5653.0148.4828.365.00
Llama-3.1-70B61.0070.3167.4769.7040.3045.00
Llama-3.2-90B58.3368.7566.2766.6738.8130.00
Qwen2.5-72B57.0062.5071.0863.6441.7910.00
Qwen2.5-14B53.6757.8171.0859.0932.8420.00
Gemma-2.0-27B49.0039.0671.0859.0931.3415.00
Gemma-2.0-9B47.3350.0050.6059.0934.3330.00
Llama-3.1-8B44.0053.1357.8345.4523.8820.00
Qwen2.5-7B43.3348.4449.4042.4238.8120.00
Mistral-Nemo43.3346.8850.6050.0032.8415.00
Mixtral-8x7B43.0053.1354.2245.4526.8710.00
Phi-3.5-mini37.0040.6348.1946.9720.900.00
Llama-3.2-3B33.0031.2549.4031.8225.370.00
Human86.6792.1990.3681.8188.0665.00

Citation

If you use this dataset, please cite the original paper:

@article{dihan2024mapeval,
  title={MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models},
  author={Dihan, Mahir Labib and Hassan, Md Tanvir and Parvez, Md Tanvir and Hasan, Md Hasebul and Alam, Md Almash and Cheema, Muhammad Aamir and Ali, Mohammed Eunus and Parvez, Md Rizwan},
  journal={arXiv preprint arXiv:2501.00316},
  year={2024}
}