Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes3k downloads1y agoHugging Face02code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes1.6k downloads2y agoHugging Face03common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes1.3k downloads1y agoHugging Face04codeparrot /github-jupyter GitHub Jupyter Dataset Dataset Description The dataset was extracted from Jupyter Notebooks on BigQuery. Licenses Each example has the license of its associated repository. There are in total 15 licenses: [ 'mit', 'apache-2.0', 'gpl-3.0', 'gpl-2.0', 'bsd-3-clause', 'agpl-3.0', 'lgpl-3.0', 'lgpl-2.1', 'bsd-2-clause', 'cc0-1.0', 'epl-1.0', 'mpl-2.0', 'unlicense', 'isc', 'artistic-2.0' ] texttext-generation100K<n<1M5 likes1.2k downloads4y agoHugging Face05LiXiang12 /github-code-fontend-lang github-code fontend code Dwonload 方式一 huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code 方式二 进入Files and versions/data直接下载zip文件 数据统计 textquestion-answering10M<n<100M2 likes775 downloads2y agoHugging Face06code-rag-bench /github-reposThe entire dump of GitHub repositories. text100K<n<1M3 likes632 downloads2y agoHugging Face07Gitnbghb /github-actionsgeospatialn<1K0 likes597 downloads27d agoHugging Face08lewtun /github-issues Dataset Card for GitHub Issues Dataset Summary GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond. Supported Tasks and Leaderboards For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.tabular1K<n<10K12 likes551 downloads5y agoHugging Face09FalconNet /GitHub-code-dialogs-1.2K-v0.1 Github Codes This is first version of dataset. All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407 text1K<n<10K1 likes438 downloads2y agoHugging Face10KoalaAI /GitHub-CC0 Public Domain GitHub Repositories Dataset This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars. The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb. The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.texttext-generation1M<n<10M6 likes372 downloads3y agoHugging Face11leongl /1c_githubtexttext-generation1M<n<10M6 likes291 downloads3y agoHugging Face12JonathanSum /github-issuestabular1K<n<10K3 likes267 downloads5y agoHugging Face13Motahar /github-issuestabulartext-retrieval1K<n<10K1 likes251 downloads4y agoHugging Face14Yatoro /github-issuestabular1K<n<10K0 likes164 downloads5y agoHugging Face15DoyyingFace /github-embeddings-doytext1K<n<10K0 likes155 downloads5y agoHugging Face16artemis13fowl /github-issuestabular1K<n<10K0 likes136 downloads5y agoHugging Face17ikumasudo /github-issuestabular1K<n<10K1 likes130 downloads5y agoHugging Face18bigscience-catalogue-data-dev /lm_code_github-eval_subsettext10K<n<100K2 likes129 downloads5y agoHugging Face19TylerHilbert /PyTorchConference2025_GithubRepos PyTorch Conference 2025 GitHub Repos I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week. textn<1K1 likes124 downloads1mo agoHugging Face20gemmozero /ai-github-ai-2026gated Ai Github Ai 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-github-ai-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-github-ai-2026.tabulartext-classificationn<1K0 likes102 downloads6d agoHugging Face21thomwolf /github-pythontext100K<n<1M10 likes92 downloads5y agoHugging Face22vchirrav /github-actions-provenance-signed-untampered GitHub Actions Provenance (Signed, Untampered) This dataset contains signed provenance metadata in JSONL format generated from successful and untampered CI/CD builds using GitHub Actions. The data represents clean, valid examples of what software artifact provenance should look like in secure, uncompromised environments. 📂 Dataset Structure Format: .jsonl files (each line is a JSON object) Source: Generated by GitHub Actions CI/CD pipelines Signature: Signed using… See the full description on the dataset page: https://huggingface.co/datasets/vchirrav/github-actions-provenance-signed-untampered.textn<1K0 likes85 downloads1y agoHugging Face23muellerzr /github-pr-history What is this dataset? This dataset is a collection of Pull Requests that contain comments from the Accelerate. It contains the full contextual comments as well as code suggestions that exist inside of a code review textn<1K0 likes77 downloads4y agoHugging Face24suolyer /pile_githubtext10K<n<100K1 likes71 downloads4y agoHugging Face25DeepNLP /Coding-Agent-Github-2025-Feb Coding Agent AI Agent Directory to Host All Coding Agent related AI Agents Web Traffic Data, Search Ranking, Community, Reviews and More. This is the Coding Agent Dataset from pypi package "coding_agent" https://pypi.org/project/coding_agent. You can use this package to download and get statistics (forks/stars/website traffic) of AI agents on website from AI Agent Marketplace AI Agent Directory (http://www.deepnlp.org/store/ai-agent) and AI Agent Search Portal… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Coding-Agent-Github-2025-Feb.textn<1K8 likes66 downloads2y agoHugging Face26alexkstern /github-code-nanochatbpe-1B github-code-nanochatbpe-1B GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 1,000,000,000 val.bin val 10,000,000 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.tabularn<1K0 likes63 downloads3mo agoHugging Face27edbeeching /github-issuesannotations_creators: other language_creators: crowdsourced languages: en-US licenses: other-my-license multilinguality: monolingual pretty_name: HuggingFace Github Issues size_categories: unknown source_datasets: original task_categories: text-classification text-retrieval task_ids: multi-class-classification multi-label-classification document-retrieval tabular1K<n<10K0 likes57 downloads5y agoHugging Face28ewhk9887 /korean_code_reviews_from_githubtext10K<n<100K1 likes56 downloads2y agoHugging Face29Agnuxo /github-source-code-dataset Github Source Code Dataset Complete source code from Agnuxo projects. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generation1K<n<10K0 likes53 downloads5mo agoHugging Face30Alexey240577 /github-trending GitHub Trending: топ репозиториев за неделю (2026-09-28) Открытые данные GitHub API: топ-10 репозиториев за 7 дней. tabularothern<1K0 likes47 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.