Team Ai
Modelpublic

Wouze/laya-ara-rag

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
3likes281downloads
Model Card

laya-ara-rag

Arabic short-list reranker — passage relevance and k≤12 ranking on `laya-multilingual`.

<p align="center"> <img src="assets/logo.jpg" alt="laya-ara-rag" width="280"> </p>

Mohammad Alkhenizan · 21 September 2026 · LinkedIn

![downloads](https://huggingface.co/Wouze/laya-ara-rag)

`Hugging Face` · `GitHub` · NLU: `laya-ara`

Highlights vs laya-multilingual

Relative lift is (fine-tune − stock) / stock. Metric is top-1 among ≤12 candidates, not corpus nDCG@10.

Arabic task*n*Base**laya-ara-rag**Relative lift
MIRACL-ar rerank (dev)28960.1530.588+284% (3.8×)
Mr.TyDi-ar rerank20000.1760.632+259% (3.6×)
MLQA-ar rerank20000.1330.582+338% (4.4×)
SadeemQuestion rerank20890.1680.871+418% (5.2×)
XPQA-ar rerank7500.2130.665+212% (3.1×)
In-domain pairwise relevance1120.8120.938+16%

Mintaka entity ranking does not improve (0.403 predecessor vs 0.388). Full tables: `RESULTS.md`. This card’s JSON: `rag_all_benches.json`. All cards: `all_cards.json`.

Abstract

laya-ara-rag is a fine-tune of `convaiinnovations/laya-multilingual` (mmBERT-base with a Laya typed-decision head, ~322M) for Arabic pairwise relevance (noul) and listwise reranking (choice). It is not generative and not a first-stage retriever. Training uses official Laya RLCD on two RTX 3090 GPUs.

The mix continues from a MIRACL-ar Wikipedia specialist, then adds in-house Arabic retrieval logs plus a MIRACL-ar train replay and a small MASSIVE-ar cap. XNLI is withheld. Evaluation is top-1 among a shortlist of k≤12.

Inference

![Open In Colab](https://colab.research.google.com/github/ASNB-Smart-Solutions/laya-ara/blob/master/examples/colab_rag.ipynb)

laya is the ConvAI Laya runtime (pip install laya==0.3.4). The Hub Use this model snippet passes trust_remote_code=True and then calls laya.load. After loading, call predict (or model.predict if you used the Hub snippet). The snippet also prints this example.

bash
pip install "laya==0.3.4"
export USE_TF=0
python
import os
import laya

agent = laya.load("Wouze/laya-ara-rag", token=os.environ.get("HF_TOKEN"))
out = agent.predict(
    {"query": "ما حكم الوضوء قبل قراءة القرآن؟"},
    {
        "passage": {
            "type": "choice",
            "instructions": "Which passage is most relevant to the query?",
            "criteria": {
                "a": "الوضوء شرط للصلاة لا للقراءة عند جمهور الفقهاء.",
                "b": "زكاة الفطر تجب على كل مسلم قبل صلاة العيد.",
                "c": "صيام عاشوراء سنة مؤكدة عند الحنابلة.",
            },
        },
        "relevant": {
            "type": "noul",
            "instructions": "Is passage A relevant to the query?",
        },
    },
)
print(out["answers"])

The example passages are illustrative. Demo: `examples/predict_rerank.py`. Scores: `rag_all_benches.json`.

Method

  • —Base. laya-multilingual (Apache-2.0). Init from the MIRACL-ar specialist, not from laya-ara.
  • —Objective. Official Laya RLCD, DDP, fp16. One epoch, 35,591 items.
  • —Public mix. MIRACL-ar train pairwise + listwise; MASSIVE-ar cap 2,000.
  • —In-house mix. Arabic retrieval logs (not redistributed): teacher-distilled pairs with a minimum score gap, listwise gold = teacher rank-1, sealed hold-out. See `DATA.md`.
  • —Protocol. Frozen JSONL. Pairwise snippets ~800 characters; listwise ~400 so k = 12 fits max_len 1024.

Public listwise rerank (k ≤ 12)

Chance ≈ 0.08–0.11. MIRACL-ar train is in the mix (dev is in-family). Other rows are transfer.

Task*n*KindBaselaya-ara**laya-ara-rag**Δ vs base
MIRACL-ar (dev)2896in-family0.1530.2100.588+284%
Mr.TyDi-ar2000transfer0.1760.2600.632+259%
SadeemQuestion2089transfer0.1680.2420.871+418%
MLQA-ar2000transfer0.1330.2140.582+338%
XPQA-ar750transfer0.2130.2810.665+212%
PublicHealthQA-ar86transfer0.1160.2440.488+321%
Mintaka-ar (entity)2203transfer0.2850.3110.388+36%

Mintaka listwise favors the Wikipedia-only predecessor (0.403). PublicHealthQA n is small.

Public pairwise relevance

One gold + one negative. Chance = 0.50.

Task*n*Baselaya-ara**laya-ara-rag**
MIRACL-ar (dev)57920.6880.6290.788
Mr.TyDi-ar40000.7550.7400.834
SadeemQuestion41780.8370.9190.941
PublicHealthQA-ar1720.7910.7730.866
Mintaka-ar44060.5370.6160.513
MLQA-ar40000.7310.7840.706
XPQA-ar15000.6230.6970.615

Listwise transfer does not imply pairwise transfer. Entity and product pairwise still favor the NLU card.

In-domain exam

Sealed hold-out on the in-house Arabic retrieval stack (queries and passages not in this repository).

Exam*n*Base**laya-ara-rag**Δ vs base
Human pairwise1120.8120.938+16%
Teacher pairwise hold4000.6400.792+24%
Teacher listwise hold2000.1100.575+423%

Teacher-hold is distillation agreement, not clicks. The 21-query human listwise set is too small to headline.

Related models

ModelRole
`laya-ara`Arabic intent, NLI, OSACT-A
laya-ara-rag (this)Arabic short-list relevance / rerank

Contact

Licensing, evaluation access, or collaboration: Mohammad Alkhenizan on LinkedIn.

Limitations

Not a first-stage retriever and not official MIRACL nDCG@10. No span extraction or token NER. Mintaka / MLQA / XPQA pairwise can drop versus laya-ara.

bibtex
@misc{alkhenizan2026layaararag,
  title        = {laya-ara-rag: Arabic short-list reranking with laya-multilingual},
  author       = {Alkhenizan, Mohammad},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Wouze/laya-ara-rag}}
}