jayansh21/codesheriff-bug-classifier
CodeSheriff Bug Classifier
A fine-tuned CodeBERT model that classifies Python code snippets into five bug categories. Built as the classification engine inside CodeSheriff — an AI system that automatically reviews GitHub pull requests.
Base model: microsoft/codebert-base · Task: 5-class sequence classification · Language: Python
Labels
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("jayansh21/codesheriff-bug-classifier")
model = AutoModelForSequenceClassification.from_pretrained("jayansh21/codesheriff-bug-classifier")
LABELS = {
0: "Clean",
1: "Null Reference Risk",
2: "Type Mismatch",
3: "Security Vulnerability",
4: "Logic Flaw"
}
code = """
def get_user(uid):
query = "SELECT * FROM users WHERE id=" + uid
return db.execute(query)
"""
inputs = tokenizer(code, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
pred = logits.argmax(dim=-1).item()
confidence = probs[0][pred].item()
print(f"{LABELS[pred]} ({confidence:.1%})")
# Security Vulnerability (99.3%)Training
Dataset: CodeSearchNet Python split with heuristic labeling, augmented with seed templates for underrepresented classes. Final training set: 4,600 balanced samples across all five classes. Stratified 80/10/10 train/val/test split.
Key hyperparameters:
Evaluation
Test set: 840 samples (stratified).
Confusion matrix:
Clean NullRef TypeMis SecVuln Logic
Actual Clean [ 394 52 1 1 2 ]
Actual NullRef [ 23 93 1 0 3 ]
Actual TypeMis [ 3 1 71 0 0 ]
Actual SecVuln [ 4 1 1 69 0 ]
Actual Logic [ 3 0 0 0 117 ]Logic Flaw and Security Vulnerability are the strongest classes — both have clear lexical patterns. Null Reference Risk is the weakest (precision 0.63) because null-risk code closely resembles clean code structurally. Most misclassifications there are false positives rather than missed bugs.
Limitations
- Python only — not trained on other languages
- Function-level input — works best on 5–50 line snippets
- Heuristic labels — training data was pattern-matched, not expert-annotated
- Not a SAST replacement — probabilistic classifier, not a sound static analysis tool
Links
- GitHub: jayansh21/CodeSheriff
- Live demo: huggingface.co/spaces/jayansh21/CodeSheriff
