Snugasabug/laya-finetuned-rlcd
laya-finetuned-rlcd
A fine-tuned version of the Apache-2.0 `laya` multilingual decision model, trained with RLCD (Resampling for Local Calibration Distribution) on the MIT-licensed jev-playground-rlcd-v0 decision corpus.
Read this section first. The headline numbers below are measured against the soft-teacher's argmax, not independent human ground truth, on a 4,000-row training subset. They are strong evidence of improvement but should be read as "teacher-agreement accuracy + calibration," not as a human-verified decision-quality benchmark. See Limitations for the full caveat.
What it does
laya is a non-autoregressive System-1 decision model. Give it a state (an email, ticket, or JSON blob) and typed questions, and it returns calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. It never generates text, so there's nothing to parse and nothing to hallucinate.
This checkpoint is fine-tuned for better accuracy and, more importantly, better calibration — a model whose reported probabilities are honest.
Results
Independent held-out set: 75,311 items (27,496 choice / 27,498 score / 20,317 noul), constructed as the full jev gold corpus minus the 4,000 training rows, so it is genuinely out-of-distribution. Both models run through the identical preprocess/forward path, so the comparison is apples-to-apples.
Two takeaways:
- Accuracy roughly doubled (0.403 → 0.959). The
scoretask is the most dramatic: base scored below random (0.173 on a 4-way task); the fine-tuned model scores 0.988. - Calibration collapsed toward ideal. RLCD fits one temperature per question type. Base needs 5.6–10.0 (its logits are badly overconfident). The fine-tuned model needs ~1.07–1.20 — essentially no correction. A temperature near 1.0 is the RLCD target: the model's own probabilities are trustworthy.
Brier (a proper scoring rule; lower is better) improves ~13–16× across all types.
Loading this checkpoint
This is a laya-format checkpoint, not a standard HF transformers model. It is loaded via the `laya` Python package, not via AutoModel.from_pretrained(). Upload the full laya checkpoint structure below.pip install layafrom laya.agent import Agent
# Load the fine-tuned checkpoint directly from a local path, or clone the repo first:
a = Agent(model_id_or_path="path/to/laya-finetuned-rlcd")
q = {"type": "choice",
"instructions": "Classify\n\"Please cancel my subscription.",
"criteria": ["keep", "cancellation"]}
out = a.predict(state="x", questions={"q": q})
# out["answers"]["q"]["probabilities"] -> {"keep": ..., "cancellation": ...}questions must be a dict {qid: {type, instructions, criteria}}; the output container is answers (not results).
Files in this repo
Training recipe
Reproduce with the laya finetune script (/opt/laya-git/research/scripts/finetune_single_device.py or its checkpointed --resume variant).
Limitations
- Accuracy is teacher-agreement, not ground truth. The target is the soft teacher's argmax — agreement with training labels, not an independent human check. "Perfect on the teacher" is strong evidence but not a guarantee of correctness on reality.
- Subset, not full corpus. This checkpoint was trained on 4,000 jev rows. The full corpus (~31,500 rows) is available; a full-corpus run is the natural next step.
- Slight residual overconfidence. Fine-tuned temperatures are ~1.07–1.20, not exactly 1.0 — the model nudges slightly overconfident, highest on the noul task (1.20). Not a problem, but stated honestly.
Attribution
- Base model: `convaiinnovations/laya` (Apache-2.0) — this is a derivative work; base license retained.
- Data: `soyrsoyr/jev-playground-rlcd-v0` (MIT).
