pngwn/github-issue-classifier-distilbert
039
DistilBERT GitHub Issue Classifier (bug / feature / question / support)
Full fine-tune of distilbert-base-uncased (66M params): lr 2e-5, batch 16, 4 epochs, weight decay 0.01, 10% warmup, title+body concatenated and truncated to 256 tokens (the truncation used by the NLBSE'24 competition winners).
Training data: pngwn/github-issues-4class — 1,997 balanced issues (bug/feature/question from NLBSE'24 + support class sourced from maintainer-assigned GitHub labels).
Test-set results (1,997 balanced examples)
Trails the smaller SetFit MiniLM model (pngwn/github-issue-classifier-setfit-minilm, 22M params, 0.807 macro-F1) on the same data — consistent with the NLBSE literature that contrastive SetFit few-shot recipes outperform straight fine-tunes at this scale.
Usage
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("pngwn/github-issue-classifier-distilbert")
model = AutoModelForSequenceClassification.from_pretrained("pngwn/github-issue-classifier-distilbert")
inputs = tok("App crashes when opening settings", return_tensors="pt")
preds = model(**inputs).logits.argmax(dim=-1)Labels (id → name): 0 bug, 1 feature, 2 question, 3 support.
