Team Ai
Modelpublic

pngwn/github-issue-classifier-distilbert

sourceHugging Facemitupdated 1d agoView on Hugging Face
0likes39downloads
Model Card

DistilBERT GitHub Issue Classifier (bug / feature / question / support)

Full fine-tune of distilbert-base-uncased (66M params): lr 2e-5, batch 16, 4 epochs, weight decay 0.01, 10% warmup, title+body concatenated and truncated to 256 tokens (the truncation used by the NLBSE'24 competition winners).

Training data: pngwn/github-issues-4class — 1,997 balanced issues (bug/feature/question from NLBSE'24 + support class sourced from maintainer-assigned GitHub labels).

Test-set results (1,997 balanced examples)

metricvalue
macro-F1 (4-class)0.794
accuracy0.794
macro-F1 on bug/feature/question subset0.779
F1 bug0.787
F1 feature0.788
F1 question0.710
F1 support0.891

Trails the smaller SetFit MiniLM model (pngwn/github-issue-classifier-setfit-minilm, 22M params, 0.807 macro-F1) on the same data — consistent with the NLBSE literature that contrastive SetFit few-shot recipes outperform straight fine-tunes at this scale.

Usage

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer

tok = AutoTokenizer.from_pretrained("pngwn/github-issue-classifier-distilbert")
model = AutoModelForSequenceClassification.from_pretrained("pngwn/github-issue-classifier-distilbert")
inputs = tok("App crashes when opening settings", return_tensors="pt")
preds = model(**inputs).logits.argmax(dim=-1)

Labels (id → name): 0 bug, 1 feature, 2 question, 3 support.