Team Ai
Modelpublic

followsci/bert-ai-text-detector

sourceHugging Facemitupdated 11mo agoView on Hugging Face
7likes1.2kdownloads
Model Card

BERT-based AI-Generated Academic Text Detector

A high-accuracy BERT model for detecting AI-generated academic text with 99.57% accuracy on paragraph-level samples.

Online Demo

๐ŸŒ Try the model online: https://followsci.com/ai-detection

Free web interface with real-time detection, no installation or API key required.

Model Details

Model Description

  • โ€”Model Type: BERT-base-uncased fine-tuned for binary text classification
  • โ€”Architecture: BERT-base-uncased (110M parameters)
  • โ€”Task: Binary classification (Human-written vs AI-generated text)
  • โ€”Input: Academic text paragraphs (up to 512 tokens)
  • โ€”Output: Binary label (0 = Human-written, 1 = AI-generated) with confidence scores

Training Information

  • โ€”Training Samples: 1,487,400 paragraph-level samples
  • โ€”Validation Samples: 185,930 paragraph-level samples
  • โ€”Test Samples: 185,930 paragraph-level samples
  • โ€”Total Dataset: 1,859,260 paragraphs
  • โ€”Training Data:
  • โ€”Human-written: Academic papers from arXiv
  • โ€”AI-generated: Text generated by various large language models (GPT, Claude, etc.)

Performance

Test Set Results

MetricValue
Accuracy99.57%
F1-Score99.58%
Precision99.23%
Recall99.94%
False Positive Rate0.82%
False Negative Rate0.06%

Confusion Matrix (Test Set)

Predicted: HumanPredicted: AI
Actual: Human89,740 (TN)740 (FP)
Actual: AI60 (FN)95,390 (TP)

Inference Speed: ~20,900 samples/second on RTX 3090 (batch size 64)

Usage

Quick Start

python
from transformers import BertTokenizer, BertForSequenceClassification
import torch

# Load model and tokenizer
model_name = "followsci/bert-ai-text-detector"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(model_name)
model.eval()

# Detect AI text
text = "Your academic paragraph here..."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)

with torch.no_grad():
    outputs = model(**inputs)
    probs = torch.nn.functional.softmax(outputs.logits, dim=-1)
    ai_prob = probs[0][1].item() * 100
    human_prob = probs[0][0].item() * 100
    
    print(f"AI-generated probability: {ai_prob:.1f}%")
    print(f"Human-written probability: {human_prob:.1f}%")
    
    if ai_prob > 50:
        print("Prediction: AI-generated")
    else:
        print("Prediction: Human-written")

Batch Processing

python
texts = [
    "First paragraph...",
    "Second paragraph...",
    # ... more texts
]

inputs = tokenizer(
    texts,
    return_tensors="pt",
    truncation=True,
    max_length=512,
    padding=True
)

with torch.no_grad():
    outputs = model(**inputs)
    probs = torch.nn.functional.softmax(outputs.logits, dim=-1)
    
    for i, prob in enumerate(probs):
        ai_prob = prob[1].item() * 100
        print(f"Text {i+1}: AI probability = {ai_prob:.1f}%")

Using with Transformers Pipeline

python
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="followsci/bert-ai-text-detector",
    tokenizer="followsci/bert-ai-text-detector"
)

result = classifier("Your text here...")
print(result)

Training Details

Training Configuration

  • โ€”Base Model: bert-base-uncased
  • โ€”Batch Size: 64
  • โ€”Learning Rate: 5e-5 (with linear warmup)
  • โ€”Warmup Steps: 5,000
  • โ€”Max Sequence Length: 512
  • โ€”Optimizer: AdamW
  • โ€”Epochs: 3
  • โ€”Training Time: ~11 hours (on RTX 3090)

Dataset Distribution

SplitTotal SamplesHuman (Label 0)AI (Label 1)
Train1,487,400723,780 (48.7%)763,620 (51.3%)
Validation185,93090,470 (48.7%)95,460 (51.3%)
Test185,93090,480 (48.7%)95,450 (51.3%)

Limitations

  1. 1.Domain Specificity: The model is trained primarily on academic text. Performance may degrade on:
  2. 2.Casual text or social media content
  3. 3.Technical documentation
  4. 4.Creative writing
  1. 1.Binary Classification: The model only distinguishes between "human" and "AI" text, without:
  2. 2.Identifying which AI model generated the text
  3. 3.Providing confidence intervals
  4. 4.Detecting partially AI-assisted text
  1. 1.Paragraph-Level Detection: The model is optimized for paragraph-level samples:
  2. 2.Performance on sentence-level or full-document level may vary
  3. 3.Best results achieved with structured academic paragraphs
  1. 1.False Positives: Approximately 0.82% false positive rate means some human-written text may be flagged as AI-generated.

Ethical Considerations

  • โ€”Use Case: This model is intended as a tool for academic integrity and research purposes
  • โ€”Bias: The model may reflect biases present in the training data
  • โ€”Misuse: Should not be used as the sole criterion for academic misconduct decisions
  • โ€”Transparency: Results should be interpreted with context and domain expertise

License

This model is licensed under the MIT License.

Contact

  • โ€”Email: raffoduanedonnenfeld@gmail.com

<p align="center"> Made with โค๏ธ for Academic Integrity </p>