datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard
Open Telco Leaderboard Scores
Benchmark scores for 84 models across 7 telecom-domain benchmarks, sourced from the MWC leaderboard.
This dataset publishes scores only (no energy metrics).
Files
leaderboard_scores.csv: Flat table for the dataset viewer.
leaderboard_scores.json: Structured JSON with per-model benchmark scores and standard errors.
Schema (leaderboard_scores.csv)
Core columns:
model — Model name
provider — Model provider (e.g. OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/leaderboard.RAG-QA-Leaderboard
Dataset Description
This collection includes 6 widely-used datasets for open-domain question answering and retrieval evaluation:
2WikiMultihopQA, HotpotQA,Musique,PopQA,TrivialQA,PubMedQA
Our evaluation code is at https://github.com/AQ-MedAI/RagQALeaderboard.
Leaderboard
Overall Performance of Different Models on Various Tasks:
Model
AVG
Multi-hop
Single-hop
Medical Domain
DeepSeekR1-0528
79.5
80
92.4
66
GPT-4.1-2025-04-14
78.8
81.6
92.8
62… See the full description on the dataset page: https://huggingface.co/datasets/AQ-MedAI/RAG-QA-Leaderboard.science_leaderboard_submissionThis dataset contains the results used for Science Leaderboard
