Team Ai
Modelpublic

FrameByFrame/programming-language-identification-100plus-lite

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes30downloads
README.md148 linesDownload Raw Back to root
1---2license: apache-2.03language:4- multilingual5tags:6- programming-language-identification7- code8- byte-level9- lite10pipeline_tag: text-classification11metrics:12- f113- accuracy14---15 16# programming-language-identification-100plus-lite17 18Byte-level programming-language identification across **107 languages**.19**2.35M parameters**, no20tokenizer, ships at **~9 MB fp32 / ~4.5 MB bf16**.21 22**[Open PyTorch Notebook](https://huggingface.co/FrameByFrame/programming-language-identification-100plus-lite/blob/main/lite_pytorch_demo.ipynb)** · **[Open ONNX Notebook](https://huggingface.co/FrameByFrame/programming-language-identification-100plus-lite/blob/main/lite_onnx_demo.ipynb)** — Download and run in Colab or Jupyter.23 24The architecture is `ByteHybrid` (3 × Conv1D → 1 × bidirectional attention with25RoPE → masked mean-pool → classifier head, with a 4096-bucket trigram-hash26embedding), vendored from27[PleIAs/CommonLingua](https://huggingface.co/PleIAs/CommonLingua) (Apache-2.0)28and trained from scratch on Rosetta Code + The Stack v1 across 107 canonical29programming languages.30 31## Comparison with `philomath-1209/programming-language-identification`32 333,057 test rows over the **26 labels** philomath supports. ONNX,34`CPUExecutionProvider`, batch 64.35 36| model | params | accuracy | macro F1 | weighted F1 | speed |37|---|---:|---:|---:|---:|---:|38| **programming-language-identification-100plus-lite** (ONNX) | 2.35 M | 0.9094 | **0.9410** | **0.9361** | **2.37×** |39| philomath-1209/programming-language-identification (ONNX) | 84 M | 0.8449 | 0.8445 | 0.8467 | 1.00× |40 41 42## Files43 44```45model.pt              fp32 PyTorch checkpoint (CommonLingua format)46model.bf16.pt         bf16 sidecar checkpoint (smaller, same accuracy in eval)47lang2idx.json         107-label index48training_metadata.json  hyperparameters and dataset stats49training_history.json   per-epoch loss / val_acc / val_macro_f150onnx/51  model.onnx          ONNX export (opset 20, dynamic batch)52  model.onnx.data     external weights blob53  lang2idx.json       (mirror)54  onnx_metadata.json  parity report vs PyTorch55```56 57## Quick start — PyTorch58 59```python60import torch, numpy as np, sys61sys.path.append("path/to/code-language-id/src")62from code_language_id.byte_hybrid import ByteHybrid, CONFIGS63 64ckpt = torch.load("model.pt", map_location="cpu", weights_only=False)65model = ByteHybrid(num_classes=ckpt["num_classes"], max_len=ckpt["max_len"],66                   **CONFIGS[ckpt["config"]]).eval()67model.load_state_dict(ckpt["model_state_dict"])68idx2lang = {v: k for k, v in ckpt["lang2idx"].items()}69 70def encode(texts, max_len=ckpt["max_len"]):71    out = np.full((len(texts), max_len), 256, dtype=np.int64)72    for i, t in enumerate(texts):73        b = t.encode("utf-8", errors="replace")[:max_len]74        out[i, :len(b)] = np.frombuffer(b, dtype=np.uint8)75    return torch.from_numpy(out)76 77with torch.no_grad():78    logits = model(encode(["def hello():\n    print('hi')"]))79print(idx2lang[int(logits.argmax(-1))])   # -> Python80```81 82## Quick start — ONNX Runtime83 84```python85import onnxruntime as ort, numpy as np, json86 87sess = ort.InferenceSession("onnx/model.onnx", providers=["CPUExecutionProvider"])88lang2idx = json.load(open("onnx/lang2idx.json"))89idx2lang = {v: k for k, v in lang2idx.items()}90MAX_LEN = 102391 92def encode(texts, max_len=MAX_LEN):93    out = np.full((len(texts), max_len), 256, dtype=np.int64)94    for i, t in enumerate(texts):95        b = t.encode("utf-8", errors="replace")[:max_len]96        out[i, :len(b)] = np.frombuffer(b, dtype=np.uint8)97    return out98 99logits = sess.run(None, {"byte_ids": encode(["fn main() {}"])})[0]100print(idx2lang[int(logits.argmax(-1))])   # -> Rust101```102 103## Training summary104 105- **Data**: Rosetta Code (`cakiki/rosetta-code`) + The Stack v1106  (`bigcode/the-stack`), task-split to prevent leakage.107  72,549 / 9,495 / 8,880 rows (train / val / test) across 107 canonical labels.108- **Snippets**: variable-window (64–1023 bytes) UTF-8.109- **Optimizer**: AdamW (β=0.9, 0.95, weight decay 0.01) + cosine-with-warmup,110  peak LR 3e-3, 5 % warmup, gradient clipping 1.0.111- **Schedule**: 30 epochs, bf16 autocast, batch 128 (effective 128 with112  gradient clipping; SDPA fused attention).113- **Best val macro F1**: 0.9085 @ epoch 26 (early stopped).114 115See `training_metadata.json` for the full hyperparameter dump.116 117## Citation118 119If you use this model, please cite:120 121```bibtex122@misc{mariappan2026codelangidlite,123  author    = {Mariappan, Vijayachandran},124  title     = {programming-language-identification-100plus-lite: Byte-level Programming Language Identification across 107 Languages},125  year      = {2026},126  publisher = {Hugging Face},127  url       = {https://huggingface.co/FrameByFrame/programming-language-identification-100plus-lite}128}129```130 131Upstream architecture:132 133```bibtex134@misc{commonlingua,135  author    = {{PleIAs}},136  title     = {CommonLingua: Byte-level Language Identification for 334 Languages},137  year      = {2026},138  publisher = {Hugging Face},139  url       = {https://huggingface.co/PleIAs/CommonLingua}140}141```142 143## License & attribution144 145Apache-2.0. Architecture and reference inference code derive from146**PleIAs/CommonLingua** (Apache-2.0). Trained weights and dataset curation are147original to this repository.148