Team Ai
Modelpublic

rasbt/ai-text-detector-distilbert-lora

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes10downloads
Model Card

LoRA-Tuned DistilBERT AI-Text Detector

This repository contains a LoRA adapter for classifying human-written and AI-generated text with DistilBERT. It was trained on `rasbt/human-vs-ai-50k`. Human-written text has label 0 and AI-generated text has label 1.

The adapter uses rank 8 and targets the query and value projection layers. It uses a maximum sequence length of 512 tokens and temperature scaling during inference. The recorded best validation accuracy was 99.65%.

 

Download and use

bash
hf download rasbt/ai-text-detector-distilbert-lora \
  --local-dir models/ai-text-detector-distilbert-lora
python
import json
from pathlib import Path

import torch
from peft import AutoPeftModelForSequenceClassification
from transformers import AutoTokenizer


model_dir = Path("models/ai-text-detector-distilbert-lora")
metadata = json.loads(
    (model_dir / "detector-config.json").read_text(encoding="utf-8")
)
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoPeftModelForSequenceClassification.from_pretrained(model_dir)
model.eval()

text = "Paste the text to classify here."
inputs = tokenizer(
    text,
    truncation=True,
    max_length=metadata["max_length"],
    return_tensors="pt",
)

with torch.inference_mode():
    logits = model(**inputs).logits / metadata["temperature"]
    probabilities = logits.float().softmax(dim=-1)

ai_index = metadata["label_mapping"]["ai"]
ai_probability = probabilities[0, ai_index].item()
print({"score": round(100 * ai_probability, 4)})

 

Test-set confusion matrix

[image]

detector-config.json contains the adapter, calibration, and training metadata. The recommended inference implementation is provided in the `rasbt/ai-detector` repository.

 

Related models

 

Limitations

Performance may change for text from generators, domains, languages, and editing workflows not represented in the training set. Short or partly AI-assisted text may also be harder to classify. The score should not be treated as definitive evidence that a person did or did not write a text.