Team Ai
Datasetpublic

helloadhavan/github_issues

GitHub Pull Request Bug–Fix Dataset Kaggle url A curated, high-signal dataset of real-world software bugs and fixes collected from 25 popular open-source GitHub repositories.Each entry corresponds to a single pull request (PR) and pairs contextual metadata with the exact code changes (unified diffs) that fixed the bug. This dataset is designed for: Automated program repair Bug-fix patch generation LLM-based code and debugging agents Empirical software engineering research… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/github_issues.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
4likes133downloads
README.md222 linesDownload Raw Back to root
1---2license: mit3configs:4- config_name: default5  data_files:6  - split: train7    path: data/train-*8  - split: eval9    path: data/eval-*10  - split: test11    path: data/test-*12dataset_info:13  features:14  - name: repo15    dtype: string16  - name: fix_commit17    dtype: string18  - name: buggy_commit19    dtype: string20  - name: message21    dtype: string22  - name: files23    list:24    - name: path25      dtype: string26    - name: patch27      dtype: string28    - name: additions29      dtype: int6430    - name: deletions31      dtype: int6432    - name: language33      dtype: string34  - name: timestamp35    dtype: timestamp[s]36  splits:37  - name: train38    num_bytes: 156164063939    num_examples: 11509640  - name: eval41    num_bytes: 2905408142    num_examples: 300043  - name: test44    num_bytes: 2905408145    num_examples: 300046  download_size: 54962936347  dataset_size: 161974880148task_categories:49- text-generation50- summarization51language:52- en53tags:54- code55pretty_name: Github issues dataset56size_categories:57- 100K<n<1M58---59 60# GitHub Pull Request Bug–Fix Dataset61 62[Kaggle url](https://www.kaggle.com/datasets/adhavanyuvaraj/github-issues)63 64A **curated, high-signal dataset of real-world software bugs and fixes** collected from **25 popular open-source GitHub repositories**.  65Each entry corresponds to a **single pull request (PR)** and pairs contextual metadata with the **exact code changes (unified diffs)** that fixed the bug.66 67This dataset is designed for:68 69- **Automated program repair**70- **Bug-fix patch generation**71- **LLM-based code and debugging agents**72- **Empirical software engineering research**73 74---75 76## How to use77 78install datasets python library:79 80```bash81pip install datasets82```83here is a copy paste example84```python85from datasets import load_dataset86 87# Load all splits88dataset = load_dataset("helloadhavan/github_issues")89 90print(dataset)91# pick the train split92 93example = dataset["train"][0]94 95# Inspect a single example96 97print("Repository:", example["repo"])98print("Buggy commit:", example["buggy_commit"])99print("Fix commit:", example["fix_commit"])100print("Message:", example["message"])101print("Timestamp:", example["timestamp"])102 103print("\nModified files:")104for f in example["files"]:105    print("-", f["path"], f["language"])106 107# Filter examples by programming language108 109def contains_assembly_file(example):110    return any(f["language"] == "Assembly" for f in example["files"])111 112python_fixes = dataset["train"].filter(contains_assembly_file)113 114print("Assembly-related fixes:", len(python_fixes))115 116```117 118 119## Data collection methodology120 121Data was collected from **GitHub repositories** by identifying commit pairs that represent122a **bug-introducing version** and its corresponding **fix commit**.123 124The dataset was constructed and post-processed to ensure high signal and usability:125 126- Only commits representing **bug fixes or correctness changes** were included127- Each example explicitly links a **buggy commit** to the corresponding **fix commit**128- Repository metadata is preserved for traceability129- Code changes are stored as **unified diffs at the file level**130- Commits that only perform refactoring, formatting, or non-functional changes were excluded131- Entries without meaningful code changes were filtered out132 133Each dataset row represents **one bug–fix commit pair**, rather than a pull request.134 135---136 137## Dataset schema138 139Each entry in the dataset follows the schema below:140 141```json142{143  "repo": "owner/repository",144  "buggy_commit": "abcdef123456...",145  "fix_commit": "fedcba654321...",146  "message": "Commit message describing the fix",147  "timestamp": "YYYY-MM-DDTHH:MM:SSZ",148  "files": [149    {150      "path": "path/to/file.ext",151      "patch": "unified diff representing the fix",152      "additions": 10,153      "deletions": 2,154      "language": "Programming language inferred from file extension"155    }156  ]157}158```159 160| Field               | Description                                           |161| ------------------- | ----------------------------------------------------- |162| `repo`              | GitHub repository containing the fix                  |163| `buggy_commit`      | Commit introducing or containing the bug              |164| `fix_commit`        | Commit that fixes the bug                             |165| `message`           | Commit message associated with the fix                |166| `timestamp`         | Timestamp of the fix commit (ISO 8601 format)         |167| `files`             | List of files modified by the fix                     |168| `files[].path`      | Path to the modified file                             |169| `files[].patch`     | Unified diff containing the code changes              |170| `files[].additions` | Number of lines added                                 |171| `files[].deletions` | Number of lines removed                               |172| `files[].language`  | Programming language inferred from the file extension |173 174 175## Supported languages176 177The dataset contains fixes across multiple programming languages, including (but not limited to):178 179* JavaScript / TypeScript180* C / C++181* Python182* Rust183* Go184* Java185* Objective-C / Objective-C++ (rare)186* Assembly (very rare. only 638 samples)187 188Language distribution varies by repository.189 190## Intended use cases191 192This dataset is well-suited for:193 194* Training models to generate patches from real pull request context195* Studying bug-fix patterns across large codebases196* Building autonomous debugging or repair agents197* Research in program repair, code synthesis, and software maintenance198 199It is not intended for:200 201* Pull request classification or triage202* Sentiment analysis203 204## Limitations205 206The dataset reflects real-world noise from GitHub pull requests207Buggy commit identification is heuristic and may be imperfect208Some fixes involve refactoring or design changes rather than minimal patches209No guarantee that fixes represent optimal or best-practice solutions210 211<blockquote212  style="213    background: #fff7cc;214    border-left: 5px solid #ffad00;215    padding: 12px 16px;216    color: #5c4b00;217    font-style: italic;218    border-radius: 4px;219  "220>221  <strong style="color:rgba(57, 0, 0, 1)">Note:</strong> Due to a bug in the scraper code, 121k samples were collected instead of the planned 50k.222</blockquote>