emperorfutures/github-top-code1
GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code… See the full description on the dataset page: https://huggingface.co/datasets/emperorfutures/github-top-code1.
0134
1---2license: mit3task_categories:4 - text-generation5language:6 - code7tags:8 - code9 - github10 - source-code11 - trending-developers12 - software-engineering13size_categories:14 - 1M<n<10M15---16 17# GitHub Top Developer Source Code18 19A curated dataset of 1.3M+ source code files from **GitHub's top ranked developers (2015-2025)**.20 21This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers22 23## Dataset Summary24 25- **1.3M+ source code files** from repositories across ~4,700 unique developers26- **80+ programming languages** included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)27- **Source code only** — config files (JSON, YAML, TOML, etc.) and documentation (Markdown, TXT) are excluded28- **Permissive licenses only** (MIT, Apache-2.0, BSD, ISC, etc.)29- **Rich metadata** per file: repo stars, description, primary language, developer company affiliation30 31 3233 34## Schema35 36Each row represents a single source file:37 38| Column | Type | Description |39|--------|------|-------------|40| `file_path` | string | Path within the repo (e.g. `src/main.py`) |41| `file_language` | string | Language detected from file extension (e.g. `Python`, `JavaScript`) |42| `content` | string | Raw source code (UTF-8) |43| `repo_name` | string | Full repository name (`owner/repo`) |44| `repo_stars` | int64 | GitHub star count at time of collection |45| `repo_description` | string | Repository description |46| `repo_primary_language` | string | GitHub-detected primary language of the repository |47| `developer_username` | string | GitHub username |48| `developer_name` | string | Developer display name |49| `developer_company` | string | Company affiliation |50 51**Note on language columns:** `file_language` is determined per-file from the file extension (e.g. a `.py` file is always `Python`). `repo_primary_language` is GitHub's auto-detected primary language for the entire repository. These may differ — for example, a C header file (`.h` → `C/C++ Header`) in a repo that GitHub classifies as `Python`.52 53## Splits54 55| Split | Description |56|-------|-------------|57| `train` | ~90% of repos — for training |58| `test` | ~5% of repos — for evaluation |59| `validation` | ~5% of repos — for hyperparameter tuning |60 61Splits are assigned **by repository** (deterministic hash), so no repo appears in multiple splits. This prevents data leakage from files in the same project.62 63## Usage64 65```python66from datasets import load_dataset67 68# Load a specific split69train = load_dataset("ronantakizawa/github-top-code", split="train")70test = load_dataset("ronantakizawa/github-top-code", split="test")71 72# Filter by language73python_files = train.filter(lambda x: x["file_language"] == "Python")74 75# Filter by stars76popular = train.filter(lambda x: x["repo_stars"] > 1000)77 78# Get files from a specific developer79dev_files = train.filter(lambda x: x["developer_username"] == "torvalds")80```81 