atmaneayoub/jev-ar-bench
Jev-AR Bench Arabic intent-routing benchmarks for MSA, Emirati, Saudi and code-switched Arabic 4,974 held-out routing queries · 60 routes · 4 language slices · plus an unseen-domain RAG test Two benchmarks for routing Arabic messages to the right destination. Each is split evenly across four language slices: MSA, Emirati, Saudi and mixed Arabic–English. They are the evaluation sets of Jev-AR. Benchmark Config Questions Routes What it measures… See the full description on the dataset page: https://huggingface.co/datasets/atmaneayoub/jev-ar-bench.
<div align="center">
Jev-AR Bench
Arabic intent-routing benchmarks for MSA, Emirati, Saudi and code-switched Arabic
  
4,974 held-out routing queries · 60 routes · 4 language slices · plus an unseen-domain RAG test
</div>
Two benchmarks for routing Arabic messages to the right destination. Each is split evenly across four language slices: MSA, Emirati, Saudi and mixed Arabic–English. They are the evaluation sets of Jev-AR.
from datasets import load_dataset
gov = load_dataset("atmaneayoub/jev-ar-bench", "gov_test", split="test")
health = load_dataset("atmaneayoub/jev-ar-bench", "health_rag", split="test")Leaderboard
Gulf-Arabic routing: accuracy in %, with 95% confidence intervals for Jev-AR.
Unseen-domain routing: correct decisions out of 40 (20 questions × an English and an Arabic route list).
Add your model. Open a discussion with your model, the protocol you used and your scores (per-question predictions if you can), and it will be added to the leaderboard.
Evaluate your own router
The full-60 protocol on gov_test, in a dozen lines:
import json
from datasets import load_dataset
from huggingface_hub import hf_hub_download
bench = "atmaneayoub/jev-ar-bench"
test = load_dataset(bench, "gov_test", split="test")
spec = json.load(open(hf_hub_download(bench, "gov_test/intents.json", repo_type="dataset"), encoding="utf-8"))
options = [r["name_ar"] for r in spec["routes"]] # full-60: canonical Arabic names, in this order
to_id = {r["name_ar"]: r["id"] for r in spec["routes"]}
def my_router(text, options, instruction):
"""Return one of `options` for `text`. Replace this with your system."""
...
correct = sum(to_id[my_router(q["text"], options, spec["instruction"])] == q["label"] for q in test)
print(f"full-60 accuracy: {correct / len(test):.1%}")With Jev-AR as my_router (router.route(text, options, instruction)["route"]), this loop gives 83.5% (4,151 / 4,974) in about 50 s on one GPU. The leaderboard's 83.4% (4,149) comes from the batched evaluation harness; the two-question difference is bf16 numerics from padding.
Gulf-Arabic routing (gov_test)
4,974 questions over a 58-intent UAE + KSA government-services taxonomy (12 categories), plus 2 escape options, other_government and not_government. They were written by DeepSeek-V4.1-Flash from realistic scenarios and frozen before any model was trained. SHA-256 of gov_test/test.jsonl: af7333663398a753aeac0c1aa78ce3b050ea95a0e880c64310f769aea098a1a6.
gov_test/intents.json lists every option with its Arabic and English names, plus the instruction.
- Protocols:
- full-60: all 60 options with canonical Arabic names, in the order of
intents.json. - sub-20: the gold option, the 2 escapes, and random options up to 20.
Unseen-domain routing (health_rag)
One health-insurance RAG app with 6 routes: FAQ, escalate to a human, benefits, exclusions, network check and nearest in-network provider. It has 20 questions, 5 per slice, and each has exactly one correct route. No insurer's app, and no routes like these, were in Jev-AR's training data. The government data does include one public service about mandatory health insurance (676 training questions), so the topic is not new, but the app and its routes are. It was written with Claude and reviewed by the author.
- Protocol: route every question over the English route list and over the Arabic list, in listed order. The score is correct answers out of 40. Also report how many answers change when the list is reversed.
- Files:
health_rag/health_rag_v1.jsonis the canonical file: routes with English and Arabic names and descriptions, the instruction, and the questions. Its content SHA-256 (parsed, key-sorted JSON) isd1c8fb26b512460b06576014865a949d2e0717095df7971983068236ec700340.questions.jsonlandroutes.jsonlhold the same content, flattened for the viewer. - Runner:
benchmark/run_health_rag.pyon GitHub.
Quality and limitations
- Synthetic: every question was written by an LLM, and none was validated by native speakers.
- Label check: an independent pass by DeepSeek (the generator's own family) agrees with the gold label on 94.2% of
gov_test. Some adversarial labels are debatable. For a cleaner subset, keep the 4,686 questions where the labeller agrees. - Size:
health_ragis small. With 20 questions, a 2-point difference out of 40 is not significant. - The benchmarks are not affiliated with any government or company.
Licence
CC BY 4.0 (`LICENSE`). gov_test was generated with DeepSeek-V4.1-Flash (MIT; its API terms allow using the outputs), and health_rag with Claude. Native-speaker reviews of the dialect slices are welcome.
Citation
@misc{jevarbench2026,
title = {Jev-AR Bench: Arabic Intent-Routing Benchmarks},
author = {atmaneayoub},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/atmaneayoub/jev-ar-bench}}
}