Team Ai
Modelpublic

cavi-ai/laya-MLX-8bit

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
2likes286downloads
Model Card

laya-MLX-8bit

MLX conversion of convaiinnovations/laya for Apple Silicon: the root (English) checkpoint and the multilingual and typed-decisions checkpoints.

  • —Converted from the base model at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.
  • —Not affiliated with or endorsed by convaiinnovations.
  • —Converted by Sasan Sotoodehfar, CAVI AI (https://cavi-ai.xyz).

Contents

PathCheckpointEncoderContext (`max_len` / `head_max_len`)Size
/laya (English)ModernBERT-large, 28 layers512 / 192524 MB
multilingual/laya-multilingualmmBERT-base, 22 layers, 256k vocabulary1024 / 256569 MB
typed-decisions/laya-typed-decisionsModernBERT-large, 28 layers1024 / 256524 MB
  • —Each folder holds model.safetensors, config.json, tokenizer.json, tokenizer_config.json.
  • —Quantized to 8-bit (affine, group size 32): the linear layers of the encoder and of the decision head.
  • —Kept in float16: token embeddings, norms, type embedding, option scorer, act head.
  • —config.json carries the decision settings and the shipped temperatures (temperature, temperature_by_options).
  • —laya/: MLX model code (ModernBERT encoder from mlx-embeddings, decision head, request rendering, calibrated answers). mlx-embeddings 0.1.0 does not include this model.
  • —8-bit only: at 4 bits the multilingual checkpoint changed 3 of 25 reference answers.

Requirements

  • —Apple Silicon Mac.
  • —Python 3.12.
  • —pip install mlx-embeddings==0.1.0

Usage

python
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("cavi-ai/laya-MLX-8bit")
sys.path.insert(0, path)
from laya import load, predict

model, tokenizer = load(path)  # or f"{path}/multilingual", f"{path}/typed-decisions"
state = {"document": "I was charged twice for my subscription this month. Please refund the duplicate charge."}
questions = {
    "department": {"type": "choice", "instructions": "Which team should handle this ticket?",
                   "criteria": {"billing": "payments, invoices, refunds", "technical": "bugs, outages, errors", "sales": "new purchases and upgrades"}},
    "urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["not urgent", "somewhat urgent", "very urgent"]},
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
}
print(predict(model, tokenizer, state, questions)["answers"])
  • —state: text or a JSON object.
  • —Question types: choice (criteria as a list or an object of option: description), score (ordered list of levels), noul (optional true / false texts).
  • —Answers: choice with probabilities and confidence; score (expected level) with probabilities; noul (probability of true).
  • —All questions of a request run in one forward pass.

Measured results

Hardware: Apple M5 Max.

Checkroot`multilingual``typed-decisions`
Reference questions17 (English)25 (7 languages)17 (English)
Same answer as the PyTorch fp32 reference17/1725/2517/17
Mean / max probability change vs the reference0.0009 / 0.0110.0026 / 0.0210.0008 / 0.003
Usage example above: department, refund probabilitybilling 0.95, 0.89billing 1.00, 0.99billing 0.82, 0.73
  • —Reference: transformers ModernBertModel plus torch.nn decision-head layers, float32, on CPU, loaded from the original checkpoints; tokenization through AutoTokenizer.
  • —Port check: the MLX code in float32 on CPU matches the reference within 2.5e-5 on every option logit, with identical input ids.
  • —One request with three questions answers in 12–46 ms after loading.

License

  • —Model weights, tokenizer, and configuration files: Apache-2.0, inherited from the base model. See LICENSE.
  • —Code in laya/: MIT. See laya/LICENSE.

Links

  • —Base model: https://huggingface.co/convaiinnovations/laya
  • —MLX port source: https://github.com/cavi-ai/mlx-agent