Team Ai
Modelpublic

FrameByFrame/programming-language-identification-100plus

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes37downloads
README.md127 linesDownload Raw Back to root
1---2license: apache-2.03library_name: transformers4pipeline_tag: text-classification5tags:6- text-classification7- code8- programming-language-identification9- language-detection10- modernbert11base_model: answerdotai/ModernBERT-base12datasets:13- cakiki/rosetta-code14- bigcode/the-stack15metrics:16- accuracy17- f118---19 20# Programming Language Identification (100+ languages)21 22A ModernBERT classifier that identifies the programming language of a code23snippet across **107 languages**.24 25## Inference26 27### PyTorch28 29```python30import torch31from transformers import AutoModelForSequenceClassification, AutoTokenizer32 33model_id = "FrameByFrame/programming-language-identification-100plus"34tokenizer = AutoTokenizer.from_pretrained(model_id)35model = AutoModelForSequenceClassification.from_pretrained(36    model_id,37    attn_implementation="eager",38    torch_dtype=torch.bfloat16,39).eval()40 41code = "def greet(name: str) -> None:\n    print(f'hello, {name}')"42inputs = tokenizer(code, return_tensors="pt", truncation=True, max_length=512)43with torch.no_grad():44    logits = model(**inputs).logits45print(model.config.id2label[int(logits.argmax(-1))])  # -> "Python"46```47 48### Batch49 50```python51snippets = [py_code, rust_code, go_code]  # list of strings52inputs = tokenizer(53    snippets, return_tensors="pt", padding=True, truncation=True, max_length=51254)55with torch.no_grad():56    logits = model(**inputs).logits57for i, pred in enumerate(logits.argmax(-1).tolist()):58    print(snippets[i][:40].splitlines()[0], "→", model.config.id2label[pred])59```60 61### ONNX Runtime62 63An ONNX export lives in `onnx/`. Use it for CPU or GPU inference without64pulling PyTorch — handy for non-Python consumers and edge deployments.65 66```python67from optimum.onnxruntime import ORTModelForSequenceClassification68from transformers import AutoTokenizer69 70model_id = "FrameByFrame/programming-language-identification-100plus"71tokenizer = AutoTokenizer.from_pretrained(model_id)72ort_model = ORTModelForSequenceClassification.from_pretrained(73    model_id, subfolder="onnx"74)75 76inputs = tokenizer(code, return_tensors="pt", truncation=True, max_length=512)77logits = ort_model(**inputs).logits78print(ort_model.config.id2label[int(logits.argmax(-1))])79```80 81**[Open Inference Notebook](https://huggingface.co/FrameByFrame/programming-language-identification-100plus/blob/main/inference_examples.ipynb)** — download and run in Colab or Jupyter.82 83## Evaluation84 85Held-out validation split (9,495 rows, 107 labels):86 87| metric | value |88|---|---|89| macro F1 | **0.9206** |90| accuracy | 0.9306 |91 92 93Wins on every shared label. Largest gaps: ARM Assembly +0.354, Erlang +0.270,94COBOL +0.216, Pascal +0.206, Fortran +0.193, Mathematica/Wolfram +0.173.95 96## Supported languages (107)97 98ABAP, APL, ARM Assembly, ATS, Ada, ActionScript, AppleScript, AutoHotkey,99AutoIt, Awk, BASIC, BQN, Batchfile, Befunge, C, C#, C++, COBOL, Ceylon,100Clojure, CoffeeScript, ColdFusion, Common Lisp, Component Pascal, Crystal, D,101Dart, E, Eiffel, Elixir, Emacs Lisp, Erlang, Euphoria, F#, Factor, Fantom,102Forth, Fortran, FreeBASIC, GAP, Go, Groovy, Haskell, Haxe, IDL, Io, J, Java,103JavaScript, Julia, Kotlin, LabVIEW, LFE, Lasso, Logtalk, Lua, M, M4, MATLAB,104MAXScript, Mathematica/Wolfram Language, Mercury, Modula-2, Modula-3, Nemerle,105NewLisp, Nim, OCaml, Objective-C, Oz, PHP, Pascal, Perl, Pike, PicoLisp,106PowerShell, Processing, Prolog, PureBasic, Python, QuickBASIC, R, REXX, Raku,107Racket, Rebol, Red, Ring, Ruby, Rust, SAS, Scala, Scheme, Scilab, Smalltalk,108Standard ML, Stata, Swift, Tcl, V, VBA, VBScript, Vala, Visual Basic .NET,109Wren, Zig, jq110 111## Training data112 11391,209 code samples across 107 languages, drawn from Rosetta Code114(`cakiki/rosetta-code`) and The Stack v1 (`bigcode/the-stack`). Labels were115independently verified by an LLM judge, and a small set of high-confidence116mislabels between mainstream languages was removed.117 118Splits are grouped by task to prevent task-level leakage:11972,549 / 9,495 / 8,880 rows (train / val / test).120 121## Limitations122 123- Only the first **512 characters** of each input are used — longer files are124  truncated before classification.125- The classifier is purely content-based. If you have file extensions, treat126  them as a strong prior in a production pipeline.127