eulogik/pico-type
<div align="center">
pico-type π
A tiny byte-level multi-head content classifier β ~1.5M params, ~9MB single-file ONNX (FP32), ~18ms CPU inference.
Classifies any content from raw bytes: coarse type Β· modality Β· subtype Β· code language Β· text language Β· file MIME Β· risk flags
 ![Python]()  ![ONNX]()   
</div>
β¨ Features
- No tokenizer β operates directly on raw UTF-8 bytes (supports all languages, no preprocessing)
- 7 heads, one forward pass β coarse type, modality, subtype, code language, text language, file MIME, risk flags
- 4 Matryoshka tiers β tiny (16d) β small (64d) β base (192d) β pro (576d) β same trunk, accuracy scales with dim
- ~9MB single-file ONNX (FP32) β deploy on edge devices, serverless, browser (WebAssembly/ONNX Runtime Web)
- ~18ms inference on CPU via ONNX Runtime
- CLI, Python API, Gradio Space, MCP server β ready to use
π‘οΈ ARTH V2 β Risk++ (new)
[Try it live β pick the `arth` model](https://huggingface.co/spaces/eulogik/pico-type) Β· Model files (arth_full_base.onnx, risk_thresholds.json, temperatures.json, arth_final_composite.pt)
Frozen-trunk composite (ft3 student + retrained 14-label Risk++ head) β audit recall up 5Γ with zero regressions on every shipped gate:
Formerly-dead labels now detected: jwt / sshkey / email **8/8**, password 0.875, phone 0.750. Artifact: **`arthfull_base.onnx` (11.66 MB, opset 18)** β auto-verified (torch-vs-ORT err 1.1e-05, parity 60/60, semantic 300/300, risk flags 11/11); INT8 4.28 MB experimental.
Known limitations (measured): piissn 0.125, apikey 0.625, jailbreak 0.50 on embedded phrasing; 3/11 hand-picked benign probes still flag (incl. the accepted print('hello world')βsql); long-document signals dilute in fixed-window pooling; general-classification gates (AG/SST-2/Enron) stay at chance β this is a byte-pattern risk flagger, not a general text classifier.
ARTH quickstart (ONNX)
import json, math
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
sess = ort.InferenceSession(hf_hub_download("eulogik/pico-type", "arth_full_base.onnx"))
thrs = json.load(open(hf_hub_download("eulogik/pico-type", "risk_thresholds.json")))
LABELS = ["api_key","jwt","ssh_key","password","email","phone","prompt_injection",
"jailbreak","pii_ssn","pii_card","secrets_aws","secrets_github",
"sql_injection","xss_payload"]
raw = open("file.txt","rb").read()[:1024]
ids = np.zeros(1024, np.int64); ids[:len(raw)] = list(raw)
mask = np.zeros(1024, bool); mask[:len(raw)] = True
(logits,) = sess.run(["riskpp_logits"], {
"input_ids": ids[None,:], "attention_mask": mask[None,:],
"opt_embs": np.zeros((1,4,96), np.float32),
"struct_feats": np.zeros((1,14), np.float32),
"act_stats": np.zeros((1,4), np.float32)})
probs = [1/(1+math.exp(-x)) for x in logits[0]]
print({l: round(p,4) for l,p in zip(LABELS, probs) if p >= thrs[l]}) # fired flagsARTH quickstart (live Space API)
from gradio_client import Client
c = Client("eulogik/pico-type")
out = c.predict("aws_access_key_id = AKIAIOSFODNN7EXAMPLE", "arth", api_name="/handle_classify")
print(out[7]) # Risk++ (14) tab: probs + [FLAG] markersπ Evaluation
Overall Accuracy (v2 β trained on real data)
v0.1 baseline (synthetic-only): code_lang 3%, text_lang 19%. Real-data training in v2 improves code by 57pp and text by 79pp.
Code Language β Per-Language Accuracy
Note: Low-accuracy languages have fewer real training samples. More data will improve them.
π Quick Start
Install
pip install picotypeCLI
# Classify from stdin
echo "def hello(name):\n return f'Hi {name}'" | picotype --pretty
# Classify a file
picotype --file document.txt
# Classify clipboard content
picotype --clip
# All 4 tiers available
echo "..." | picotype --tier proPython API
from picotype import load_onnx_model, run_onnx
session = load_onnx_model("base")
result = run_onnx(session, "def hello(): pass")
print(result)
# {
# "coarse": "code",
# "code_language": "python",
# "modality": "textual",
# "confidence": 0.98,
# ...
# }MCP Server (for Claude Desktop, Cursor, etc.)
pip install picotype
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_serverThen add to your MCP config:
{
"mcpServers": {
"pico-type": {
"command": "python",
"args": ["-m", "model.pico_type.mcp_server"],
"env": { "PICOTYPE_MODEL_DIR": "./checkpoints" }
}
}
}Gradio Web UI
Try it live: huggingface.co/spaces/eulogik/pico-type
π Architecture
Bytes ββΆ ByteEmbed(256β96d) ββΆ 3ΓConv1D(k=3,5,7) ββΆ 2ΓBiAttention(RoPE) ββΆ Pool ββΆ 7ΓMatryoshka HeadsTotal parameters: 1.43M (tiny) / 1.45M (small) / 1.48M (base) / 1.56M (pro)
π§ Model Tiers
ONNX sizes are single-file FP32 exports (graph-only files are 203β206 KB).
All tiers share the same backbone; only the final linear projection layers differ. Higher-tier models use more dimensions for finer-grained classification.
π§ͺ Classification Heads
π Deployment
π Resources
- Paper β Architecture, training, and evaluation details
- Model Card β Detailed architecture and training configuration
- Walkthrough β Development log and decisions
- Architecture Plan β Original design document
π License
Apache 2.0
<div align="center"> <sub>Built with PyTorch Β· ONNX Β· Gradio Β· HuggingFace</sub> </div>
