Team Ai
Modelpublic

jayansh21/codesheriff-bug-classifier

sourceHugging Faceupdated 7mo agoView on Hugging Face
1likes16downloads
Model Card

CodeSheriff Bug Classifier

A fine-tuned CodeBERT model that classifies Python code snippets into five bug categories. Built as the classification engine inside CodeSheriff — an AI system that automatically reviews GitHub pull requests.

Base model: microsoft/codebert-base · Task: 5-class sequence classification · Language: Python


Labels

IDLabelExample
0CleanWell-formed code, no issues
1Null Reference Riskresult.fetchone().name without a None check
2Type Mismatch"Error: " + error_code where error_code is an int
3Security Vulnerability"SELECT * FROM users WHERE id = " + user_id
4Logic Flawfor i in range(len(items) + 1)

Usage

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("jayansh21/codesheriff-bug-classifier")
model = AutoModelForSequenceClassification.from_pretrained("jayansh21/codesheriff-bug-classifier")

LABELS = {
    0: "Clean",
    1: "Null Reference Risk",
    2: "Type Mismatch",
    3: "Security Vulnerability",
    4: "Logic Flaw"
}

code = """
def get_user(uid):
    query = "SELECT * FROM users WHERE id=" + uid
    return db.execute(query)
"""

inputs = tokenizer(code, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    logits = model(**inputs).logits

probs = torch.softmax(logits, dim=-1)
pred = logits.argmax(dim=-1).item()
confidence = probs[0][pred].item()

print(f"{LABELS[pred]} ({confidence:.1%})")
# Security Vulnerability (99.3%)

Training

Dataset: CodeSearchNet Python split with heuristic labeling, augmented with seed templates for underrepresented classes. Final training set: 4,600 balanced samples across all five classes. Stratified 80/10/10 train/val/test split.

Key hyperparameters:

ParameterValue
Epochs4
Effective batch size16 (8 × 2 grad accum)
Learning rate2e-5
OptimizerAdamW + linear warmup
Max token length512
Class weightingYes — balanced
HardwareNVIDIA RTX 3050 (4GB)

Evaluation

Test set: 840 samples (stratified).

ClassPrecisionRecallF1Support
Clean0.920.880.90450
Null Reference Risk0.630.780.70120
Type Mismatch0.960.950.9575
Security Vulnerability0.990.920.9575
Logic Flaw0.960.970.97120
Macro F10.890.900.89

Confusion matrix:

                 Clean  NullRef  TypeMis  SecVuln  Logic
Actual Clean   [  394      52        1        1      2  ]
Actual NullRef [   23      93        1        0      3  ]
Actual TypeMis [    3       1       71        0      0  ]
Actual SecVuln [    4       1        1       69      0  ]
Actual Logic   [    3       0        0        0    117  ]

Logic Flaw and Security Vulnerability are the strongest classes — both have clear lexical patterns. Null Reference Risk is the weakest (precision 0.63) because null-risk code closely resembles clean code structurally. Most misclassifications there are false positives rather than missed bugs.


Limitations

  • —Python only — not trained on other languages
  • —Function-level input — works best on 5–50 line snippets
  • —Heuristic labels — training data was pattern-matched, not expert-annotated
  • —Not a SAST replacement — probabilistic classifier, not a sound static analysis tool

Links