Team Ai
Datasetpublic

atmaneayoub/jev-ar-bench

Jev-AR Bench Arabic intent-routing benchmarks for MSA, Emirati, Saudi and code-switched Arabic 4,974 held-out routing queries · 60 routes · 4 language slices · plus an unseen-domain RAG test Two benchmarks for routing Arabic messages to the right destination. Each is split evenly across four language slices: MSA, Emirati, Saudi and mixed Arabic–English. They are the evaluation sets of Jev-AR. Benchmark Config Questions Routes What it measures… See the full description on the dataset page: https://huggingface.co/datasets/atmaneayoub/jev-ar-bench.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
2likes743downloads
Dataset Card

<div align="center">

Jev-AR Bench

Arabic intent-routing benchmarks for MSA, Emirati, Saudi and code-switched Arabic

![Model](https://huggingface.co/atmaneayoub/jev-ar) ![GitHub](https://github.com/atmaneayoubdev/jev-ar) ![License: CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)

4,974 held-out routing queries · 60 routes · 4 language slices · plus an unseen-domain RAG test

</div>

Two benchmarks for routing Arabic messages to the right destination. Each is split evenly across four language slices: MSA, Emirati, Saudi and mixed Arabic–English. They are the evaluation sets of Jev-AR.

BenchmarkConfigQuestionsRoutesWhat it measures
Gulf-Arabic routinggov_test4,97460UAE + KSA government-service routing, with 519 adversarial questions, 6 routes held out from training, and requests outside the list
Unseen-domain routinghealth_rag206routing over a RAG app's own route list, for a kind of app never seen in training
python
from datasets import load_dataset
gov = load_dataset("atmaneayoub/jev-ar-bench", "gov_test", split="test")
health = load_dataset("atmaneayoub/jev-ar-bench", "health_rag", split="test")

Leaderboard

Gulf-Arabic routing: accuracy in %, with 95% confidence intervals for Jev-AR.

ModelParams60 routes20 routes
Qwen3.8-27B (LLM, list in prompt)27B90.094.0
Jev 1.13 (TypeSafe, closed API)undisclosed86.291.4
Jev-AR322M83.4 <sub>[81.2–85.4]</sub>90.0 <sub>[88.4–91.4]</sub>
mmBERT-base, fine-tuned classifier307M76.579.8
multilingual-E5-base + logistic regression278M72.477.4
Laya multilingual (Jev-AR's starting point)322M42.056.4

Unseen-domain routing: correct decisions out of 40 (20 questions × an English and an Arabic route list).

ModelScoreAnswers changed when the list is reversed
Jev 1.13 (TypeSafe, closed API)400
Qwen3.8-27B400
Jev-AR380
Laya multilingual (Jev-AR's starting point)1311
Add your model. Open a discussion with your model, the protocol you used and your scores (per-question predictions if you can), and it will be added to the leaderboard.

Evaluate your own router

The full-60 protocol on gov_test, in a dozen lines:

python
import json
from datasets import load_dataset
from huggingface_hub import hf_hub_download

bench = "atmaneayoub/jev-ar-bench"
test = load_dataset(bench, "gov_test", split="test")
spec = json.load(open(hf_hub_download(bench, "gov_test/intents.json", repo_type="dataset"), encoding="utf-8"))
options = [r["name_ar"] for r in spec["routes"]]  # full-60: canonical Arabic names, in this order
to_id = {r["name_ar"]: r["id"] for r in spec["routes"]}

def my_router(text, options, instruction):
    """Return one of `options` for `text`. Replace this with your system."""
    ...

correct = sum(to_id[my_router(q["text"], options, spec["instruction"])] == q["label"] for q in test)
print(f"full-60 accuracy: {correct / len(test):.1%}")

With Jev-AR as my_router (router.route(text, options, instruction)["route"]), this loop gives 83.5% (4,151 / 4,974) in about 50 s on one GPU. The leaderboard's 83.4% (4,149) comes from the batched evaluation harness; the two-question difference is bf16 numerics from padding.

Gulf-Arabic routing (gov_test)

4,974 questions over a 58-intent UAE + KSA government-services taxonomy (12 categories), plus 2 escape options, other_government and not_government. They were written by DeepSeek-V4.1-Flash from realistic scenarios and frozen before any model was trained. SHA-256 of gov_test/test.jsonl: af7333663398a753aeac0c1aa78ce3b050ea95a0e880c64310f769aea098a1a6.

FieldMeaning
text, labelthe message and its correct intent id
category, slice, countrythe service category, the language slice, and uae / ksa
kindintent (3,612), adversarial (519), not_gov (430) or other_gov (413)
adv_typedistractor, negation, tworequests, veryshort, othercountry, nearmissprivate or politechit_chat
held_outtrue for the 6 intents that Jev-AR never saw in training
situationthe scenario the question was written from

gov_test/intents.json lists every option with its Arabic and English names, plus the instruction.

  • —Protocols:
  • —full-60: all 60 options with canonical Arabic names, in the order of intents.json.
  • —sub-20: the gold option, the 2 escapes, and random options up to 20.

Unseen-domain routing (health_rag)

One health-insurance RAG app with 6 routes: FAQ, escalate to a human, benefits, exclusions, network check and nearest in-network provider. It has 20 questions, 5 per slice, and each has exactly one correct route. No insurer's app, and no routes like these, were in Jev-AR's training data. The government data does include one public service about mandatory health insurance (676 training questions), so the topic is not new, but the app and its routes are. It was written with Claude and reviewed by the author.

  • —Protocol: route every question over the English route list and over the Arabic list, in listed order. The score is correct answers out of 40. Also report how many answers change when the list is reversed.
  • —Files: health_rag/health_rag_v1.json is the canonical file: routes with English and Arabic names and descriptions, the instruction, and the questions. Its content SHA-256 (parsed, key-sorted JSON) is d1c8fb26b512460b06576014865a949d2e0717095df7971983068236ec700340. questions.jsonl and routes.jsonl hold the same content, flattened for the viewer.
  • —Runner: benchmark/run_health_rag.py on GitHub.

Quality and limitations

  • —Synthetic: every question was written by an LLM, and none was validated by native speakers.
  • —Label check: an independent pass by DeepSeek (the generator's own family) agrees with the gold label on 94.2% of gov_test. Some adversarial labels are debatable. For a cleaner subset, keep the 4,686 questions where the labeller agrees.
  • —Size: health_rag is small. With 20 questions, a 2-point difference out of 40 is not significant.
  • —The benchmarks are not affiliated with any government or company.

Licence

CC BY 4.0 (`LICENSE`). gov_test was generated with DeepSeek-V4.1-Flash (MIT; its API terms allow using the outputs), and health_rag with Claude. Native-speaker reviews of the dialect slices are welcome.

Citation

bibtex
@misc{jevarbench2026,
  title        = {Jev-AR Bench: Arabic Intent-Routing Benchmarks},
  author       = {atmaneayoub},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/atmaneayoub/jev-ar-bench}}
}