Team Ai
Datasetpublic

emperorfutures/github-top-code1

GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code… See the full description on the dataset page: https://huggingface.co/datasets/emperorfutures/github-top-code1.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes134downloads
README.md81 linesDownload Raw Back to root
1---2license: mit3task_categories:4  - text-generation5language:6  - code7tags:8  - code9  - github10  - source-code11  - trending-developers12  - software-engineering13size_categories:14  - 1M<n<10M15---16 17# GitHub Top Developer Source Code18 19A curated dataset of 1.3M+ source code files from **GitHub's top ranked developers (2015-2025)**.20 21This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers22 23## Dataset Summary24 25- **1.3M+ source code files** from repositories across ~4,700 unique developers26- **80+ programming languages** included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)27- **Source code only** — config files (JSON, YAML, TOML, etc.) and documentation (Markdown, TXT) are excluded28- **Permissive licenses only** (MIT, Apache-2.0, BSD, ISC, etc.)29- **Rich metadata** per file: repo stars, description, primary language, developer company affiliation30 31 32![Screenshot 2026-02-23 at 10.41.38 AM](https://cdn-uploads.huggingface.co/production/uploads/65a752167bfcb01564e6276c/WJVtEjZijsz8zW3KT0TU4.png)33 34## Schema35 36Each row represents a single source file:37 38| Column | Type | Description |39|--------|------|-------------|40| `file_path` | string | Path within the repo (e.g. `src/main.py`) |41| `file_language` | string | Language detected from file extension (e.g. `Python`, `JavaScript`) |42| `content` | string | Raw source code (UTF-8) |43| `repo_name` | string | Full repository name (`owner/repo`) |44| `repo_stars` | int64 | GitHub star count at time of collection |45| `repo_description` | string | Repository description |46| `repo_primary_language` | string | GitHub-detected primary language of the repository |47| `developer_username` | string | GitHub username |48| `developer_name` | string | Developer display name |49| `developer_company` | string | Company affiliation |50 51**Note on language columns:** `file_language` is determined per-file from the file extension (e.g. a `.py` file is always `Python`). `repo_primary_language` is GitHub's auto-detected primary language for the entire repository. These may differ — for example, a C header file (`.h` → `C/C++ Header`) in a repo that GitHub classifies as `Python`.52 53## Splits54 55| Split | Description |56|-------|-------------|57| `train` | ~90% of repos — for training |58| `test` | ~5% of repos — for evaluation |59| `validation` | ~5% of repos — for hyperparameter tuning |60 61Splits are assigned **by repository** (deterministic hash), so no repo appears in multiple splits. This prevents data leakage from files in the same project.62 63## Usage64 65```python66from datasets import load_dataset67 68# Load a specific split69train = load_dataset("ronantakizawa/github-top-code", split="train")70test = load_dataset("ronantakizawa/github-top-code", split="test")71 72# Filter by language73python_files = train.filter(lambda x: x["file_language"] == "Python")74 75# Filter by stars76popular = train.filter(lambda x: x["repo_stars"] > 1000)77 78# Get files from a specific developer79dev_files = train.filter(lambda x: x["developer_username"] == "torvalds")80```81