Wouze/laya-ara-rag
laya-ara-rag
Arabic short-list reranker — passage relevance and k≤12 ranking on `laya-multilingual`.
<p align="center"> <img src="assets/logo.jpg" alt="laya-ara-rag" width="280"> </p>
Mohammad Alkhenizan · 21 September 2026 · LinkedIn

`Hugging Face` · `GitHub` · NLU: `laya-ara`
Highlights vs laya-multilingual
Relative lift is (fine-tune − stock) / stock. Metric is top-1 among ≤12 candidates, not corpus nDCG@10.
Mintaka entity ranking does not improve (0.403 predecessor vs 0.388). Full tables: `RESULTS.md`. This card’s JSON: `rag_all_benches.json`. All cards: `all_cards.json`.
Abstract
laya-ara-rag is a fine-tune of `convaiinnovations/laya-multilingual` (mmBERT-base with a Laya typed-decision head, ~322M) for Arabic pairwise relevance (noul) and listwise reranking (choice). It is not generative and not a first-stage retriever. Training uses official Laya RLCD on two RTX 3090 GPUs.
The mix continues from a MIRACL-ar Wikipedia specialist, then adds in-house Arabic retrieval logs plus a MIRACL-ar train replay and a small MASSIVE-ar cap. XNLI is withheld. Evaluation is top-1 among a shortlist of k≤12.
Inference

laya is the ConvAI Laya runtime (pip install laya==0.3.4). The Hub Use this model snippet passes trust_remote_code=True and then calls laya.load. After loading, call predict (or model.predict if you used the Hub snippet). The snippet also prints this example.
pip install "laya==0.3.4"
export USE_TF=0import os
import laya
agent = laya.load("Wouze/laya-ara-rag", token=os.environ.get("HF_TOKEN"))
out = agent.predict(
{"query": "ما حكم الوضوء قبل قراءة القرآن؟"},
{
"passage": {
"type": "choice",
"instructions": "Which passage is most relevant to the query?",
"criteria": {
"a": "الوضوء شرط للصلاة لا للقراءة عند جمهور الفقهاء.",
"b": "زكاة الفطر تجب على كل مسلم قبل صلاة العيد.",
"c": "صيام عاشوراء سنة مؤكدة عند الحنابلة.",
},
},
"relevant": {
"type": "noul",
"instructions": "Is passage A relevant to the query?",
},
},
)
print(out["answers"])The example passages are illustrative. Demo: `examples/predict_rerank.py`. Scores: `rag_all_benches.json`.
Method
- Base.
laya-multilingual(Apache-2.0). Init from the MIRACL-ar specialist, not from laya-ara. - Objective. Official Laya RLCD, DDP, fp16. One epoch, 35,591 items.
- Public mix. MIRACL-ar train pairwise + listwise; MASSIVE-ar cap 2,000.
- In-house mix. Arabic retrieval logs (not redistributed): teacher-distilled pairs with a minimum score gap, listwise gold = teacher rank-1, sealed hold-out. See `DATA.md`.
- Protocol. Frozen JSONL. Pairwise snippets ~800 characters; listwise ~400 so k = 12 fits
max_len1024.
Public listwise rerank (k ≤ 12)
Chance ≈ 0.08–0.11. MIRACL-ar train is in the mix (dev is in-family). Other rows are transfer.
Mintaka listwise favors the Wikipedia-only predecessor (0.403). PublicHealthQA n is small.
Public pairwise relevance
One gold + one negative. Chance = 0.50.
Listwise transfer does not imply pairwise transfer. Entity and product pairwise still favor the NLU card.
In-domain exam
Sealed hold-out on the in-house Arabic retrieval stack (queries and passages not in this repository).
Teacher-hold is distillation agreement, not clicks. The 21-query human listwise set is too small to headline.
Related models
Contact
Licensing, evaluation access, or collaboration: Mohammad Alkhenizan on LinkedIn.
Limitations
Not a first-stage retriever and not official MIRACL nDCG@10. No span extraction or token NER. Mintaka / MLQA / XPQA pairwise can drop versus laya-ara.
@misc{alkhenizan2026layaararag,
title = {laya-ara-rag: Arabic short-list reranking with laya-multilingual},
author = {Alkhenizan, Mohammad},
year = {2026},
howpublished = {\url{https://huggingface.co/Wouze/laya-ara-rag}}
}