Team Ai
Modelpublic

Snugasabug/laya-finetuned-rlcd

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes48downloads
Model Card

laya-finetuned-rlcd

A fine-tuned version of the Apache-2.0 `laya` multilingual decision model, trained with RLCD (Resampling for Local Calibration Distribution) on the MIT-licensed jev-playground-rlcd-v0 decision corpus.

Read this section first. The headline numbers below are measured against the soft-teacher's argmax, not independent human ground truth, on a 4,000-row training subset. They are strong evidence of improvement but should be read as "teacher-agreement accuracy + calibration," not as a human-verified decision-quality benchmark. See Limitations for the full caveat.

What it does

laya is a non-autoregressive System-1 decision model. Give it a state (an email, ticket, or JSON blob) and typed questions, and it returns calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. It never generates text, so there's nothing to parse and nothing to hallucinate.

This checkpoint is fine-tuned for better accuracy and, more importantly, better calibration — a model whose reported probabilities are honest.

Results

Independent held-out set: 75,311 items (27,496 choice / 27,498 score / 20,317 noul), constructed as the full jev gold corpus minus the 4,000 training rows, so it is genuinely out-of-distribution. Both models run through the identical preprocess/forward path, so the comparison is apples-to-apples.

TypenBASEFINETUNEDΔBASE raw ECEFT raw ECEBASE BrierFT BrierBASE tempFT temp
choice27,4960.4510.984+0.5330.3020.13716,7511,3786.701.07
score27,4980.1730.988+0.8150.5050.13322,6901,35310.01.10
noul20,3170.6470.885+0.2380.2200.0377,3811,4485.581.20
overall75,3110.4030.959+0.556

Two takeaways:

  1. 1.Accuracy roughly doubled (0.403 → 0.959). The score task is the most dramatic: base scored below random (0.173 on a 4-way task); the fine-tuned model scores 0.988.
  2. 2.Calibration collapsed toward ideal. RLCD fits one temperature per question type. Base needs 5.6–10.0 (its logits are badly overconfident). The fine-tuned model needs ~1.07–1.20 — essentially no correction. A temperature near 1.0 is the RLCD target: the model's own probabilities are trustworthy.

Brier (a proper scoring rule; lower is better) improves ~13–16× across all types.

Loading this checkpoint

This is a laya-format checkpoint, not a standard HF transformers model. It is loaded via the `laya` Python package, not via AutoModel.from_pretrained(). Upload the full laya checkpoint structure below.
bash
pip install laya
python
from laya.agent import Agent

# Load the fine-tuned checkpoint directly from a local path, or clone the repo first:
a = Agent(model_id_or_path="path/to/laya-finetuned-rlcd")

q = {"type": "choice",
     "instructions": "Classify\n\"Please cancel my subscription.",
     "criteria": ["keep", "cancellation"]}
out = a.predict(state="x", questions={"q": q})
# out["answers"]["q"]["probabilities"] -> {"keep": ..., "cancellation": ...}

questions must be a dict {qid: {type, instructions, criteria}}; the output container is answers (not results).

Files in this repo

FilePurpose
model.safetensorsFine-tuned weights (643 MB)
rl_agent_config.jsonModel config, incl. fitted temperatures [1.071, 1.034, 1.045]
encoder/mmBERT/ModernBert encoder
tokenizer/Tokenizer

Training recipe

Baseconvaiinnovations/laya (Apache-2.0), multilingual
Datajev-playground-rlcd-v0 (MIT); 4,000-row decision subset
ObjectiveRLCD — policy-gradient loss over noisy logit projections (GRPO-style, 4 samples/item, exploration noise annealed 0.4→0.1), rewarded by proper scoring rules (spherical 0.75, ranked-probability 1.0); plus full-weight soft cross-entropy
Epochs4
OptimizerAdamW, cosine schedule, WD 0.01
LRsencoder 2.5e-5, head 1e-4
Batchmicrobatch 8, groupsize 4, maxtokensper_batch 4096
Lengthsmaxlen 1024, headmax_len 256
DeviceCPU

Reproduce with the laya finetune script (/opt/laya-git/research/scripts/finetune_single_device.py or its checkpointed --resume variant).

Limitations

  1. 1.Accuracy is teacher-agreement, not ground truth. The target is the soft teacher's argmax — agreement with training labels, not an independent human check. "Perfect on the teacher" is strong evidence but not a guarantee of correctness on reality.
  2. 2.Subset, not full corpus. This checkpoint was trained on 4,000 jev rows. The full corpus (~31,500 rows) is available; a full-corpus run is the natural next step.
  3. 3.Slight residual overconfidence. Fine-tuned temperatures are ~1.07–1.20, not exactly 1.0 — the model nudges slightly overconfident, highest on the noul task (1.20). Not a problem, but stated honestly.

Attribution