Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmanPriyanshu /random-small-github-repositories random-small-github-repositories A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. Contents seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash) repos-zipped/ — one .zip per repo, named {repo_hash}.zip unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.tabulartext-generation1K<n<10K0 likes317 downloads6mo agoHugging Face02AmanPriyanshu /random-python-github-repositories random-python-github-repositories A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files. Contents repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash) repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.tabulartext-generation1K<n<10K0 likes315 downloads6mo agoHugging Face03shailja /Verilog_GitHub VeriGen Dataset Summary The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery. For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces). Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation Point of Contact: contact@shailja Languages:… See the full description on the dataset page: https://huggingface.co/datasets/shailja/Verilog_GitHub.text100K<n<1M32 likes282 downloads3y agoHugging Face04aurelium /github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table. It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and totals of the amounts of various events on those repositories in that time period. This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing, so I figured I would save the next person to try this some effort. tabular100M<n<1B7 likes252 downloads2y agoHugging Face05ronantakizawa /github-top-projects GitHub Trending Projects (2013-2025) A comprehensive dataset of 423,098 GitHub trending repository entries spanning 12+ years (August 2013 - November 2025), scraped from Wayback Machine snapshots of GitHub's trending page. 🎯 Dataset Overview This dataset captures the evolution of GitHub's trending repositories over time, providing insights into: Software development trends across programming languages and domains Popular open-source projects and their trending patterns… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-projects.tabulartext-classification100K<n<1M11 likes210 downloads10mo agoHugging Face06av9ash /GitBugs GitBugs https://github.com/av9ash/gitbugs/ License and Citation This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following: @article{patil2025gitbugs, title={GitBugs: Bug Reports for Duplicate Detection, Retrieval Augmented Generation, Triage, and More}, author={Patil, Avinash}, journal={arXiv preprint arXiv:2504.09651}, year={2025} } For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/GitBugs.textsentence-similarity1K<n<10K0 likes198 downloads6mo agoHugging Face07JDhruv14 /Bhagavad-Gita_Dataset Srimad Bhagavad Gita Dataset A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks. Dataset Details Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita_Dataset.tabulartranslationn<1K61 likes170 downloads1y agoHugging Face08CooperBench /qwen35-9b-git-coop What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git). Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent (step_limit=300) Setting coop + git remote Repos 18 Pairs 211 Both-pass 5.7% (12/210… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop.tabularn<1K0 likes141 downloads4mo agoHugging Face09burak29 /git-natural-language-commands Git Natural Language Commands A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands. Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.texttext-generation1K<n<10K1 likes135 downloads26d agoHugging Face10getstarhunt /githubstats GitHub Developer Index: reachability and availability by city Aggregated statistics on 85,938 active GitHub developers in the United States, France and the United Kingdom, broken down by metropolitan area, technology and engineering role. Each row reports how many developers the group contains, what share publish an email address on their GitHub profile, and what share have flagged themselves as available for hire. Last updated: 2026-10-02. Refreshed monthly. Files… See the full description on the dataset page: https://huggingface.co/datasets/getstarhunt/githubstats.tabular1K<n<10K1 likes127 downloads7d agoHugging Face11JDhruv14 /Bhagavad-Gita-QA Bhagavad-Gita-QA-Multilingual Dataset Summary Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework. This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.tabularquestion-answering10K<n<100K5 likes120 downloads1y agoHugging Face12seniruk /git-diff_to_commit_msg Hi, I’m Seniru Epasinghe 👋 I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier. 🌐 Connect with me          There are 2 version of this dataset: git-diff_to_commit_msg - 1.5K rows huggingface link kaggle link git-diff_to_commit_msg_large - 1.75M rows huggingface link kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg.text1K<n<10K0 likes99 downloads1y agoHugging Face13h1alexbel /github-readmestabularn<1K4 likes81 downloads2y agoHugging Face14JetBrains /git_good_bench Dataset Summary GitGoodBench Lite is a subset of 900 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios). The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram. This dataset thus contains 150 samples per sample type and programming language. All data in this dataset are collected from 479 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench.tabularn<1K1 likes67 downloads11mo agoHugging Face15aswin1906 /github-advisory-2023text1K<n<10K2 likes64 downloads3y agoHugging Face16JetBrains /git_good_bench-lite Dataset Summary GitGoodBench Lite is a subset of 120 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios). The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram. This dataset thus contains 20 samples per sample type and programming language. All data in this dataset are collected from 100 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-lite.tabularn<1K1 likes63 downloads11mo agoHugging Face17JetBrains /git_good_bench-train Dataset Summary GitGoodBench Lite is a subset of 17469 samples for collecting trajectories of AI agents resolving git tasks (see Supported Scenarios) for model training purposes. We support the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit chain. All data in this dataset are collected from 816 unique, open-source GitHub repositories with permissive licenses that have >= 1000 stars, >= 5 branches, >= 10 contributors and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-train.tabular10K<n<100K2 likes57 downloads11mo agoHugging Face18hreyulog /GitHub-3Repo-7User-Opinion-Dynamics GitHub 3Repo 7User Opinion Dynamics This dataset contains monthly opinion-dynamics time series derived from three large open-source GitHub repositories: Ceph, PyTorch, and Swift. Each CSV file represents one repository and contains a repository label, one timestamp column, and seven anonymized developer trajectory columns. Files file rows anonymized developer columns ceph.csv 13 7 pytorch.csv 13 7 swift.csv 13 7 Schema Each CSV… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/GitHub-3Repo-7User-Opinion-Dynamics.tabulartabular-regressionn<1K0 likes55 downloads3mo agoHugging Face19abhishekbora09 /github_repositoriestabular1M<n<10M1 likes51 downloads3y agoHugging Face20Saptak123 /Bhagavad-Gita_Dataset Srimad Bhagavad Gita Dataset A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks. Dataset Details Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/Saptak123/Bhagavad-Gita_Dataset.tabulartranslationn<1K0 likes42 downloads8mo agoHugging Face21seniruk /git-diff_to_commit_msg_large Hi, I’m Seniru Epasinghe 👋 I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier. 🌐 Connect with me          There are 2 version of this dataset: git-diff_to_commit_msg - 1.5K rows huggingface link kaggle link git-diff_to_commit_msg_large - 1.75M rows huggingface link kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg_large.text1M<n<10M0 likes41 downloads1y agoHugging Face22ronantakizawa /github-top-developers GitHub Top Developers by Year (2015-2025) A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine. 📊 Dataset Overview Total Entries: 8,125 ranked developers Years Covered: 2015 - 2025 (11 years) Unique Developers: 4,763 Source: Derived from Wayback Machine snapshots of GitHub trending developers Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.tabulartext-classification10K<n<100K71 likes40 downloads9mo agoHugging Face23Roy229 /fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001 fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001 A curated registry of points of interest in downtown Portland, Oregon. License This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY). You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source. Contents data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.tabularn<1K0 likes36 downloads2mo agoHugging Face24deepkyu /github-as-altmetric About dataset We construct this dataset for our study, which investigates the correlation between GitHub communication metrics and citation counts, examining the potential of the metrics as an altmetric. Currently, it contains about 12,000 samples of publications, which are published by top-tier AI conferences. The citation counts and the corresponding GitHub metrics might need to be updated. We strive our best to update more conferences and keep values up-to-date.… See the full description on the dataset page: https://huggingface.co/datasets/deepkyu/github-as-altmetric.tabulartabular-regression10K<n<100K1 likes34 downloads3y agoHugging Face25rahul7star /Gita-Train Bhagavad-Gita-QA-Multilingual Dataset Summary Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework. This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/Gita-Train.tabularquestion-answering10K<n<100K0 likes34 downloads1y agoHugging Face26jason1966 /abdullahkhan70_github-tech-stack-languages-and-frameworks GitHub Tech Stack Languages & Frameworks Comprehensive Repository Data: JavaScript, Python, Go, Rust & More Dataset Info Source: Kaggle Original Size: 2.17 MB Kaggle Downloads: 62 Files: 17 Files Mirrored from Kaggle tabular10K<n<100K0 likes32 downloads6mo agoHugging Face27doanhieung /gitsum GitSum Dataset Dataset Description This dataset is a collection of data originally provided by the MDEGroup GitSum repository. Citation If you use this dataset, please cite the original repository: @misc{MDEGroup_GitSum, title={GitSum}, author={MDEGroup}, year={2023}, howpublished={\url{https://github.com/MDEGroup/GitSum}}, } textsummarization1K<n<10K0 likes29 downloads2y agoHugging Face28nasa-impact /nasa-science-github-repos NASA Science GitHub Repositories A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences. This dataset is designed to support research on information retrieval and discoverability of open-source scientific software. Licensing and Intellectual Property This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.tabulartext-retrieval1K<n<10K2 likes27 downloads6mo agoHugging Face29zhuq41 /github_fetch_huggingface_terminal_7959_0292f1a2_sales_transactions Retail Sales Transactions Transaction-level retail sales records for the North American region. textn<1K0 likes27 downloads2mo agoHugging Face30brunoziie /git_commitstextn<1K0 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.