datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.random-python-github-repositories
random-python-github-repositories
A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files.
Contents
repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash)
repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.Verilog_GitHub
VeriGen
Dataset Summary
The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/shailja/Verilog_GitHub.github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table.
It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and
totals of the amounts of various events on those repositories in that time period.
This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing,
so I figured I would save the next person to try this some effort.
github-top-projects
GitHub Trending Projects (2013-2025)
A comprehensive dataset of 423,098 GitHub trending repository entries spanning 12+ years (August 2013 - November 2025), scraped from Wayback Machine snapshots of GitHub's trending page.
🎯 Dataset Overview
This dataset captures the evolution of GitHub's trending repositories over time, providing insights into:
Software development trends across programming languages and domains
Popular open-source projects and their trending patterns… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-projects.GitBugs
GitBugs
https://github.com/av9ash/gitbugs/
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025gitbugs,
title={GitBugs: Bug Reports for Duplicate Detection, Retrieval Augmented Generation, Triage, and More},
author={Patil, Avinash},
journal={arXiv preprint arXiv:2504.09651},
year={2025}
}
For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/GitBugs.Bhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita_Dataset.qwen35-9b-git-coop
What this is
Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with
mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git).
Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit.
At a glance
Field
Value
Model
Qwen/Qwen3.5-9B
Agent
mini_swe_agent (step_limit=300)
Setting
coop + git remote
Repos
18
Pairs
211
Both-pass
5.7% (12/210… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop.git-natural-language-commands
Git Natural Language Commands
A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands.
Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.githubstats
GitHub Developer Index: reachability and availability by city
Aggregated statistics on 85,938 active GitHub developers in the United
States, France and the United Kingdom, broken down by metropolitan area,
technology and engineering role.
Each row reports how many developers the group contains, what share publish an
email address on their GitHub profile, and what share have flagged themselves
as available for hire.
Last updated: 2026-10-02. Refreshed monthly.
Files… See the full description on the dataset page: https://huggingface.co/datasets/getstarhunt/githubstats.Bhagavad-Gita-QA
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.git-diff_to_commit_msg
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg.github-readmesgit_good_bench
Dataset Summary
GitGoodBench Lite is a subset of 900 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 150 samples per sample type and programming language.
All data in this dataset are collected from 479 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench.github-advisory-2023git_good_bench-lite
Dataset Summary
GitGoodBench Lite is a subset of 120 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 20 samples per sample type and programming language.
All data in this dataset are collected from 100 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-lite.git_good_bench-train
Dataset Summary
GitGoodBench Lite is a subset of 17469 samples for collecting trajectories of AI agents resolving git tasks (see Supported Scenarios) for model training purposes.
We support the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit chain.
All data in this dataset are collected from 816 unique, open-source GitHub repositories with permissive licenses
that have >= 1000 stars, >= 5 branches, >= 10 contributors and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-train.GitHub-3Repo-7User-Opinion-Dynamics
GitHub 3Repo 7User Opinion Dynamics
This dataset contains monthly opinion-dynamics time series derived from three large open-source GitHub repositories: Ceph, PyTorch, and Swift. Each CSV file represents one repository and contains a repository label, one timestamp column, and seven anonymized developer trajectory columns.
Files
file
rows
anonymized developer columns
ceph.csv
13
7
pytorch.csv
13
7
swift.csv
13
7
Schema
Each CSV… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/GitHub-3Repo-7User-Opinion-Dynamics.github_repositoriesBhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/Saptak123/Bhagavad-Gita_Dataset.git-diff_to_commit_msg_large
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg_large.github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.github-as-altmetric
About dataset
We construct this dataset for our study, which investigates the correlation between GitHub communication metrics and citation counts, examining the potential of the metrics as an altmetric.
Currently, it contains about 12,000 samples of publications, which are published by top-tier AI conferences.
The citation counts and the corresponding GitHub metrics might need to be updated.
We strive our best to update more conferences and keep values up-to-date.… See the full description on the dataset page: https://huggingface.co/datasets/deepkyu/github-as-altmetric.Gita-Train
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/Gita-Train.abdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
gitsum
GitSum Dataset
Dataset Description
This dataset is a collection of data originally provided by the MDEGroup GitSum repository.
Citation
If you use this dataset, please cite the original repository:
@misc{MDEGroup_GitSum,
title={GitSum},
author={MDEGroup},
year={2023},
howpublished={\url{https://github.com/MDEGroup/GitSum}},
}
nasa-science-github-repos
NASA Science GitHub Repositories
A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission
Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics,
Planetary Science, Heliophysics, and Biological & Physical Sciences.
This dataset is designed to support research on information retrieval and
discoverability of open-source scientific software.
Licensing and Intellectual Property
This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.github_fetch_huggingface_terminal_7959_0292f1a2_sales_transactions
Retail Sales Transactions
Transaction-level retail sales records for the North American region.
git_commits
