Team Ai
Datasetpublic

pngwn/github-issues-4class

GitHub Issues 4-Class A balanced 4-class dataset for GitHub issue classification: bug / feature / question / support. Splits Split Rows Per class train 1,997 ~500 each (NLBSE rows failing a 15-char minimum-text filter dropped) test 1,997 ~500 each Columns text: issue title + \n\n + body, whitespace-normalized, capped at 4,000 chars (models truncate to 256 tokens) label: one of bug, feature, question, support repo: source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/pngwn/github-issues-4class.

sourceHugging Faceupdated 4d agoView on Hugging Face
0likes70downloads
Dataset Card

GitHub Issues 4-Class

A balanced 4-class dataset for GitHub issue classification: bug / feature / question / support.

Splits

SplitRowsPer class
train1,997~500 each (NLBSE rows failing a 15-char minimum-text filter dropped)
test1,997~500 each

Columns

  • —text: issue title + \n\n + body, whitespace-normalized, capped at 4,000 chars (models truncate to 256 tokens)
  • —label: one of bug, feature, question, support
  • —repo: source GitHub repository
  • —source: nlbse2024 or victor-betus

Provenance

  • —bug / feature / question: the NLBSE'24 issue-report-classification benchmark (3,000 balanced issues from facebook/react, tensorflow/tensorflow, microsoft/vscode, bitcoin/bitcoin, opencv/opencv), raw CSVs from github.com/nlbse2024/issue-report-classification. Labels are maintainer-assigned GitHub labels mapped by the benchmark's synonym table.
  • —support: issues labeled support-family (support, site-support-request, help wanted, type: support, ...) filtered from `victor-betus/github-issues-dataset` (114k issues, top-100 GitHub repos). Issues that also carry a conflicting class label (bug/enhancement/feature/question) were excluded, and all rows were deduplicated against the NLBSE benchmark by (repo, normalized title) — zero overlap.

Known caveats

  • —GitHub label conventions vary per repo; cross-project label inconsistency is the dominant error source in this task (Izadi et al., MSR 2024).
  • —The question class is historically the hardest (F1 ~0.70 in our runs); support is easiest.
  • —help wanted is treated as a support-request synonym following GitHub convention in several large repos; it can also denote "contributors wanted" in others — some label noise in the support class is possible.

Baseline results (this dataset, test split)

Modelmacro-F1 (4-class)
SetFit all-MiniLM-L6-v2 (contrastive + logistic head)0.807
DistilBERT fine-tune (lr 2e-5, 4 epochs, 256 tokens)see model card

References

  • —Colavito, Lanubile, Novielli, "Few-Shot Learning for Issue Report Classification", NLBSE 2023
  • —Kallis et al., "The NLBSE'24 Tool Competition" (issue report classification track)
  • —Tunstall et al., "Efficient Few-Shot Learning Without Prompts" (SetFit), arXiv:2209.11055