Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes8.2k downloads10mo agoHugging Face02ruediste /codeparrot-github-code-10GThis is data is derived from the Codeparrot Dataset by taking the first 10GB of text from each language, and splitting it into individual configs. This results in a download size of about 3GB per language. Sample usage: from datasets import load_dataset dataset = load_dataset("ruediste/codeparrot-github-code-10G", "java") List of Languages: languages = { 'HTML': 'html', 'Java': 'java', 'JavaScript': 'js', 'CSS': 'css', 'C#': 'cs', 'TypeScript': 'ts', "Batchfile":… See the full description on the dataset page: https://huggingface.co/datasets/ruediste/codeparrot-github-code-10G.text10M<n<100M2 likes3k downloads2y agoHugging Face03open-index /open-github OpenGitHub What is it? This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth. The archive currently… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.tabulartext-generation100K<n<1M9 likes2.7k downloads6mo agoHugging Face04ronantakizawa /github-top-code GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.texttext-generation1M<n<10M125 likes2.6k downloads8mo agoHugging Face05angie-chen55 /python-github-codetext1M<n<10M50 likes1.8k downloads4y agoHugging Face06angie-chen55 /javascript-github-codetext10M<n<100M18 likes1.7k downloads4y agoHugging Face07nick007x /github-code-2025text100M<n<1B121 likes1.5k downloads6mo agoHugging Face08ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M64 likes1.3k downloads7mo agoHugging Face09labofsahil /github-event-dataset-2019text100M<n<1B1 likes1.3k downloads2y agoHugging Face10teven /github_all_lang_filteredtext10M<n<100M2 likes1.1k downloads5y agoHugging Face11hardikg2907 /github-code-html-css-2text1M<n<10M0 likes1k downloads2y agoHugging Face12bigcode /github-commits-diff-dedup-pjjs-april Deduplicated Commits Deduplicated based on diff: content = '\n'.join(difflib.unified_diff( old_content.splitlines(keepends=True), new_content.splitlines(keepends=True), n=5 )) Parameters: Minimum ngram size: 5 MinHash ngram size: 5 MinHash threshold: 0.8 text100K<n<1M4 likes1k downloads3y agoHugging Face13jinaai /github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.image1K<n<10K0 likes956 downloads1y agoHugging Face14hardikg2907 /github-code-html-css-1text1M<n<10M2 likes852 downloads2y agoHugging Face15macrocosm-os /code-parrot-github-code GitHub Code Dataset Dataset Description The GitHub Code dataset consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in 1TB of data. The dataset was created from the public GitHub dataset on Google BiqQuery. How to use it The GitHub Code dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of datasets. You can load and iterate through the dataset with the following… See the full description on the dataset page: https://huggingface.co/datasets/macrocosm-os/code-parrot-github-code.texttext-generation100M<n<1B13 likes669 downloads2y agoHugging Face16kenhktsui /github-code-permissive-sampleSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'].It is intended to be used for training code language classifier. texttext-classification1M<n<10M0 likes669 downloads2y agoHugging Face17DELith /github-issuestabular10K<n<100K1 likes662 downloads5y agoHugging Face18CarperAI /github-diffs-dedupedtabular10M<n<100M4 likes587 downloads4y agoHugging Face19labofsahil /github-event-dataset-2023text1B<n<10B0 likes583 downloads11mo agoHugging Face20labofsahil /github-event-dataset-2021text100M<n<1B0 likes573 downloads11mo agoHugging Face21vietgpt /github Dataset Card for "github" More Information needed text10M<n<100M0 likes548 downloads3y agoHugging Face22hybridfree /github-code-2025 🚀 GitHub Code 2025: The Clean Code Manifesto A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem 🌟 The Philosophy Quality Over Quantity, Purpose Over Volume In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development. 🎯 What This Dataset Is… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.text100M<n<1B1 likes537 downloads9mo agoHugging Face23labofsahil /github-event-dataset-2013text10M<n<100M0 likes532 downloads2y agoHugging Face24labofsahil /github-event-dataset-2022text1B<n<10B0 likes443 downloads11mo agoHugging Face25Dhivyasri /github-issuestext10M<n<100M0 likes424 downloads2y agoHugging Face26labofsahil /github-event-dataset-2020text100M<n<1B0 likes410 downloads1y agoHugging Face27ibragim-bad /github-repos-metadata-40M 📊 Metadata for 40 million GitHub repositories A cleaned, analysis-ready dataset with per-repository statistics aggregated from GH Archive events: stars, forks, pull requests, open issues, visibility, language signals, and more. Column names mirror the GH Archive / GitHub API semantics where possible. GitHub repo: https://github.com/ibragim-bad/github-repos-metadata-40M Source: GH Archive (public GitHub event stream). ✅ Projects with complementary ideas GitHub Repo… See the full description on the dataset page: https://huggingface.co/datasets/ibragim-bad/github-repos-metadata-40M.tabular10M<n<100M23 likes399 downloads9mo agoHugging Face28mabujadallah /GitHub-Agentic-PR-Dataset GitHub Agentic PR Dataset A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers, complete with commits, file-level diffs, patches, bug-fix classification, review-depth metrics, and test-quality measures. The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.tabulartext-classification10M<n<100M1 likes385 downloads13d agoHugging Face29blastwind /github-code-haskell-file Dataset Card for "github-code-haskell-file" Rows: 339k Download Size: 806M This dataset is extracted from github-code-clean. Each row also contains attribute values for my personal analysis project. 12.6% (43k) of the rows have cyclomatic complexity and LOC valued at -1 because homplexity failed in parsing the row's uncommented_code. tabulartext-generation100K<n<1M1 likes363 downloads3y agoHugging Face30labofsahil /github-event-dataset-2024text1B<n<10B0 likes358 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.