aaaa47080/CryptoMind-Guard-0.6B
CryptoMind-Guard-0.6B
A small, CPU/iGPU-friendly pass / block classifier for community forum posts, built for the CryptoMind crypto & investing community. It flags scams/fraud, sexual content (including any sexual content involving minors), violence and threats, hate and harassment, self-harm, drugs/weapons, doxxing and other illegal activity — while letting ordinary investing discussion through (price talk, portfolio sharing, criticism, scam warnings and victims asking for help).
It is a LoRA fine-tune of Alibaba-AAIG/YuFeng-XGuard-Reason-0.6B (Qwen3-0.6B) and keeps the base model's output format: the first answer token is a risk-category code (sec = safe, ec = economic crimes / fraud, ma = minor abuse & exploitation, pc = pornographic content, …). You only need one forward pass and the probabilities of 29 tokens — no text generation.
中文摘要:論壇貼文「該不該擋」的小型分類模型(0.6B),微調自 YuFeng-XGuard-Reason-0.6B。詐騙、色情(含任何涉及未成年的性內容)、 暴力威脅、仇恨騷擾、自殘、毒品槍械、肉搜等要擋;一般投資討論、防詐提醒、被害求助不擋。只讀第一個 token 的機率,CPU/內顯都跑得動。 主要支援繁體中文、簡體中文、英文、俄文。
About CryptoMind
**CryptoMind** is an AI research assistant for crypto and stocks, with a community forum where investors share analysis, plus a scam-report tracker. CryptoMind-Guard was built to keep that forum useful: check every post and comment before it goes live, stop scams and abuse, and stay out of the way of normal investing talk. 👉 Try it at getcryptomind.com.
Files
The unquantized fp16 weights are not part of this release.
How it works
- Tokenize the post:
ids = tokenizer.encode(lead + post.strip())(leadis a single space, seereadout.json). - Long posts (more than 450 tokens) are scored as two pieces — the first 300 and the last 150 tokens — and the higher score wins (contact details and links are often at the end).
- Input =
prefix_ids + post_ids + suffix_idsfromreadout.json. This is token-for-token identical to the base model's chat template (apply_chat_template([{"role": "user", "content": post}], policy=None, reason_first=False)); the 299-token prefix is the same for every post, so llama.cpp's prompt cache makes repeated calls cheap. - Take the next-token distribution, keep only the 29
label_ids, renormalize, and use score = 1 − P(`sec`). The highest non-seccode is the most likely category.
Recommended policy: block at score ≥ 0.8; publish but send to human review at 0.5–0.8.
llama.cpp
llama-server -m cryptomind-guard-0.6b-q8_0.gguf -c 1024 -np 1 --port 8080 # add -ngl 99 for GPU / iGPUimport json, math, urllib.request
from tokenizers import Tokenizer
r = json.load(open("readout.json"))
tok = Tokenizer.from_file("tokenizer.json")
labels = dict(zip(r["label_ids"], r["label_codes"]))
def score(post: str) -> tuple[float, str]:
ids = tok.encode(r["lead"] + post.strip(), add_special_tokens=False).ids
n = r["first_tokens"] + r["last_tokens"]
pieces = [ids] if len(ids) <= n else [ids[: r["first_tokens"]], ids[-r["last_tokens"]:]]
best = (0.0, "sec")
for piece in pieces:
body = {"prompt": r["prefix_ids"] + piece + r["suffix_ids"], "n_predict": 1, "n_probs": 60,
"temperature": 0, "cache_prompt": True}
req = urllib.request.Request("http://127.0.0.1:8080/completion", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
top = json.loads(urllib.request.urlopen(req).read())["completion_probabilities"][0]["top_logprobs"]
p = {labels[t["id"]]: math.exp(t["logprob"]) for t in top if t["id"] in labels}
total = sum(p.values()) or 1.0
risk = {c: v / total for c, v in p.items() if c != "sec"}
s = 1 - p.get("sec", 0.0) / total
best = max(best, (s, max(risk, key=risk.get) if risk else "sec"))
return best
print(score("老師帶單穩賺不賠,每天 5% 收益,加 LINE 進 VIP 群")) # high score, "ec"
print(score("BTC 短線偏空,MACD 死叉,先觀望等回測 58000")) # low scoreonnxruntime
import json, numpy as np, onnxruntime as ort
from tokenizers import Tokenizer
r = json.load(open("readout.json"))
tok = Tokenizer.from_file("tokenizer.json")
sess = ort.InferenceSession("onnx/model.onnx", providers=["CPUExecutionProvider"])
safe = r["label_codes"].index("sec")
def score_piece(piece):
ids = np.array([r["prefix_ids"] + piece + r["suffix_ids"]], dtype=np.int64)
logits = sess.run(None, {"input_ids": ids})[0][0] # 29 risk-code logits
p = np.exp(logits - logits.max()); p /= p.sum()
return 1 - float(p[safe])Evaluation
Held-out test set of 1,290 items that were never used for training (near-duplicates of test items were also removed from training):
- Minors (harmful): 316 prompts labelled as sexual content involving minors in Nemotron-Safety-Guard v3 / PolyGuardMix, keeping only items whose label refers to the prompt itself.
- Minors (benign): 348 benign prompts that mention children or teenagers (parenting, school, child-protection, news).
- Forum (harmful / benign): 135 / 244 forum-style items — scam SMS (FGRC-SCD), CryptoMind's own scam and calibration posts, and "scary but benign" posts (scam warnings, victims asking for help, crime news, violent slang in trading talk).
- Nemotron zh (harmful / benign): 147 / 100 general Chinese safety prompts (labels are noisier).
Block rate on harmful sets, false-block rate on benign sets:
Fine-tuning mainly teaches the forum boundary: scams are blocked far more reliably, and investing talk and posts that merely mention children are blocked far less often. Q8_0 stays very close to fp16 (mean score difference 0.003, 5/1,290 decisions differ at 0.8); ONNX int8 drifts more (0.020, 23/1,290).
Latency (Apple M4, 2 CPU threads, warm prompt cache): llama.cpp Q8_0 median 388 ms, p95 678 ms per post; onnxruntime int8 median 1.09 s (no prefix cache). With a GPU / iGPU offload (-ngl 99) it is much faster.
Training
- Base:
Alibaba-AAIG/YuFeng-XGuard-Reason-0.6B@9016029, LoRA r=16 / alpha=32 on all attention and MLP projections, merged. - Data: ~27.5k training posts (block ≈ 54%), Traditional/Simplified Chinese, English and Russian. Harmful examples come only from the public datasets listed above plus CryptoMind-written scam posts; benign examples include CryptoMind UI text, investing chit-chat and hard negatives (scam warnings, victims' stories, child-protection notices).
- Objective (following the base model's first-token SFT format): benign →
sec; harmful → the matching code (fraud →ec, sexual content involving minors →ma, sexual →pc, self-harm →mh, drugs →dc, weapons →dw, doxxing →pp, hate →ac, harassment →cy, threats →ti). Harmful items with an unclear category use a binary loss (all risk codes vs.sec) so the model keeps choosing the category itself. Label smoothing 0.05. - 2 epochs on one Kaggle T4, lr 1e-4, effective batch 32; best epoch picked on validation.
Limitations
- Built for community posts; it is not tuned for chat transcripts or for judging LLM responses.
- It decides "should this post be blocked?", not "which law is broken"; the category code is only meaningful when the score is high.
- Benign posts that mention children are still blocked about 5.5% of the time at 0.8 — keep a human-review / appeal path.
- Weaker on some Russian scams, English "withdrawal locked, verify your seed phrase" scams and some indirect requests.
- Public safety datasets have noisy labels; numbers on the Nemotron set should be read with that in mind.
- No model is perfect: do not use it as the only safeguard for child safety. Report suspected child sexual exploitation to the authorities in your jurisdiction.
License
Apache-2.0 (see LICENSE and NOTICE). This is a modified version of YuFeng-XGuard-Reason-0.6B (Apache-2.0, Alibaba AAIG), itself based on Qwen3-0.6B (Apache-2.0, Alibaba Cloud). Training data attributions are listed in NOTICE (Nemotron-Safety-Guard-Dataset-v3 and PolyGuardMix are CC BY 4.0).
Citation
@article{yufeng-xguard,
title={YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models},
author={Alibaba AAIG},
journal={arXiv preprint arXiv:2601.15588},
year={2026}
}