Team Ai
Modelpublic

harvestsu/laya-multilingual-rknn-tensorrt

sourceHugging Faceapache-2.0updated 1d agoView on Hugging Face
0likes
Model Card

laya-multilingual — RK3588 and Jetson Orin

Converted artifacts and the acceptance data behind them, for running `convaiinnovations/laya-multilingual` on Rockchip RK3588 (RKNN) and NVIDIA Jetson Orin (TensorRT), entirely on-device.

Conversion recipes, measurements and the acceptance protocol live in the GitHub repository: https://github.com/suharvest/laya-on-rk3588-jetson

Files

FileBytesmd5What it is
oldship.s512.m16.fp16.nanguard.sharedmask.rk3588.rknn663,180,269cff6c915b98b0400f888d55f5679fa69RK3588 delivery, fp16
laya.s512.fp16.engine650,779,868ae0688b17b297b1474a18a97e9e04c77Jetson Orin delivery, fp16
oldship.s512.m16.nanguard.sharedmask.onnx1,290,265,0105c68449182db6dd94587736bad27ccffthe ONNX the RK3588 artifact was converted from
laya-multilingual.s512.m16.fp32.ng.onnx1,290,257,838191668c270e31fe4b94674b487bef1dcthe ONNX the Jetson engine was built from
calib/692 KB—8 calibration + 16 held-out samples at seq=512
golden/8 KB—ORT fp32 reference dump for those 16 samples

The two platforms do not share an ONNX graph. They come from separate exports and undergo different graph surgery; see ARTIFACTS.md in the GitHub repository.

Measured

seq = 512. top1 is the argmax over the model's live candidate classes on the 16 held-out samples, scored against the ORT fp32 reference.

Platformp50p95top1order_exact
RK3588 (Radxa ROCK 5T)509 ms≤598 ms16/1615/16
Jetson Orin NX18.9 ms24.8 ms16/1616/16
Jetson Orin Nano22.0 ms22.9 ms16/1616/16

Neither is bit-identical to the fp32 reference — fp16 diverges at the third decimal while preserving the ordering. RK3588 p50 requires the CPU governor pinned to performance.

Quantization

Neither platform reached usable accuracy when quantized.

  • —Jetson, explicit INT8 Q/DQ — this does engage the tensor cores (247/259 layers, i8i8→i32, 1.28–1.31× faster) but the per-tensor activation scales clip the outlier channels that carry the signal: the first quantized activation is clipped 7×. top1 5–6/16.
  • —Jetson, implicit INT8 PTQ — silently produces 0 int8 layers; the engine is bit-identical to fp16.
  • —Jetson, INT4 weight-only — 1.62× slower than fp16; sm_87 has no fused int4 kernel.
  • —RKNN, INT8 / w4a16 — logits collapse; w4a16 reports a flattering rel_l2 (7.1e-05) while the ranking breaks (6/16).

Full mechanisms, and where the ceiling actually is: docs/quantization.md.

Long context

seq = 8192 is not usable, and the cause is the model rather than the deployment: on a planted-fact probe (4-way choice, chance 25%) the fp32 reference itself scores 36% at 8192, indistinguishable from chance. 512 scores 52.5%. See docs/context-length.md.

License and attribution

Apache-2.0, inherited from the upstream model. laya and laya-multilingual are the work of their respective authors; this repository is downstream of them and not affiliated with them.

"Jev" is a product of TypeSafe AI, Inc. This repository is not affiliated with, endorsed by, or sponsored by TypeSafe AI, Inc. References to Jev are descriptive — laya implements the same System One typed-decision interface.