Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes3.3k downloads1y agoHugging Face02gittensor-model-hub /spark-hermes-rounds Spark-Hermes rounds Training data from a competition (Bittensor SN74) in which miners submit prose strategies and a validator runs one pinned agent (Hermes) and model (Qwen3.8-27B through r0055, SparkHermes-27B-v1 from era e1) against every strategy inside a sealed sandbox. The tasks are bug fixes drawn from SWE-bench/SWE-smith (MIT). Only the crowned strategy of each round is exported. SFT rows are its episodes that a withheld test suite verified fully. DPO pairs a… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/spark-hermes-rounds.tabularn<1K0 likes1.4k downloads2d agoHugging Face03common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes1.3k downloads1y agoHugging Face04codeparrot /github-jupyter GitHub Jupyter Dataset Dataset Description The dataset was extracted from Jupyter Notebooks on BigQuery. Licenses Each example has the license of its associated repository. There are in total 15 licenses: [ 'mit', 'apache-2.0', 'gpl-3.0', 'gpl-2.0', 'bsd-3-clause', 'agpl-3.0', 'lgpl-3.0', 'lgpl-2.1', 'bsd-2-clause', 'cc0-1.0', 'epl-1.0', 'mpl-2.0', 'unlicense', 'isc', 'artistic-2.0' ] texttext-generation100K<n<1M5 likes1.1k downloads4y agoHugging Face05LiXiang12 /github-code-fontend-lang github-code fontend code Dwonload 方式一 huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code 方式二 进入Files and versions/data直接下载zip文件 数据统计 textquestion-answering10M<n<100M2 likes736 downloads2y agoHugging Face06code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes666 downloads2y agoHugging Face07code-rag-bench /github-reposThe entire dump of GitHub repositories. text100K<n<1M3 likes557 downloads2y agoHugging Face08lewtun /github-issues Dataset Card for GitHub Issues Dataset Summary GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond. Supported Tasks and Leaderboards For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.tabular1K<n<10K12 likes534 downloads5y agoHugging Face09Gitnbghb /github-actionsgeospatialn<1K0 likes510 downloads1mo agoHugging Face10Yanbin99 /GITQA-Aug-Legacyimage100K<n<1M2 likes387 downloads3y agoHugging Face11rmems /git-ops-recovery-trajectories Git Ops Recovery Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.text1K<n<10K0 likes375 downloads18d agoHugging Face12KoalaAI /GitHub-CC0 Public Domain GitHub Repositories Dataset This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars. The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb. The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.texttext-generation1M<n<10M6 likes289 downloads3y agoHugging Face13leongl /1c_githubtexttext-generation1M<n<10M7 likes288 downloads3y agoHugging Face14cabbage972 /GitChameleon-2.0 GitChameleon 2.0 GitChameleon 2.0 is an AI coding benchmark comprising 328 Python-based problems conditioned on specific versions of popular libraries for scientific computing and web development. It evaluates whether AI code generation models can correctly use library APIs as they existed at a particular version — a challenging test of version-specific knowledge. Note: This is GitChameleon 2.0, a distinct and newer work from the original GitChameleon benchmark. Please do not… See the full description on the dataset page: https://huggingface.co/datasets/cabbage972/GitChameleon-2.0.texttext-generationn<1K2 likes261 downloads6mo agoHugging Face15JonathanSum /github-issuestabular1K<n<10K3 likes256 downloads5y agoHugging Face16Motahar /github-issuestabulartext-retrieval1K<n<10K1 likes244 downloads4y agoHugging Face17FalconNet /GitHub-code-dialogs-1.2K-v0.1 Github Codes This is first version of dataset. All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407 text1K<n<10K1 likes215 downloads2y agoHugging Face18Yatoro /github-issuestabular1K<n<10K0 likes152 downloads5y agoHugging Face19Tavernari /git-commit-message-dttext1K<n<10K5 likes149 downloads2y agoHugging Face20DoyyingFace /github-embeddings-doytext1K<n<10K0 likes137 downloads5y agoHugging Face21TylerHilbert /PyTorchConference2025_GithubRepos PyTorch Conference 2025 GitHub Repos I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week. textn<1K1 likes132 downloads1mo agoHugging Face22artemis13fowl /github-issuestabular1K<n<10K0 likes127 downloads5y agoHugging Face23ikumasudo /github-issuestabular1K<n<10K1 likes119 downloads5y agoHugging Face24gemmozero /ai-github-ai-2026gated Ai Github Ai 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-github-ai-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-github-ai-2026.tabulartext-classificationn<1K0 likes106 downloads10d agoHugging Face25bigscience-catalogue-data-dev /lm_code_github-eval_subsettext10K<n<100K2 likes105 downloads5y agoHugging Face26vchirrav /github-actions-provenance-signed-untampered GitHub Actions Provenance (Signed, Untampered) This dataset contains signed provenance metadata in JSONL format generated from successful and untampered CI/CD builds using GitHub Actions. The data represents clean, valid examples of what software artifact provenance should look like in secure, uncompromised environments. 📂 Dataset Structure Format: .jsonl files (each line is a JSON object) Source: Generated by GitHub Actions CI/CD pipelines Signature: Signed using… See the full description on the dataset page: https://huggingface.co/datasets/vchirrav/github-actions-provenance-signed-untampered.textn<1K0 likes86 downloads1y agoHugging Face27thomwolf /github-pythontext100K<n<1M10 likes84 downloads5y agoHugging Face28the-homeless-god /git-history-mcq-ru git-history-mcq-ru 805 вопросов с вариантами ответа по истории трёх открытых репозиториев (digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс 8 672 ответа пяти моделей и 4 878 разборов этих ответов. Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю. Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.tabularmultiple-choice10K<n<100K0 likes78 downloads1mo agoHugging Face29muellerzr /github-pr-history What is this dataset? This dataset is a collection of Pull Requests that contain comments from the Accelerate. It contains the full contextual comments as well as code suggestions that exist inside of a code review textn<1K0 likes69 downloads4y agoHugging Face30alexkstern /github-code-nanochatbpe-1B github-code-nanochatbpe-1B GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 1,000,000,000 val.bin val 10,000,000 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.tabularn<1K0 likes66 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.