datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spark-hermes-rounds
Spark-Hermes rounds
Training data from a competition (Bittensor SN74) in which miners submit prose strategies and a validator
runs one pinned agent (Hermes) and model (Qwen3.8-27B through r0055, SparkHermes-27B-v1 from era e1) against
every strategy inside a sealed sandbox. The tasks
are bug fixes drawn from SWE-bench/SWE-smith (MIT).
Only the crowned strategy of each round is exported. SFT rows are its episodes that a withheld test suite
verified fully. DPO pairs a… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/spark-hermes-rounds.github-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.github-actionsgithub-issuesgithub-issuesgithub-issuesgithub-issuesgithub-issuesai-github-ai-2026
Ai Github Ai 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-github-ai-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-github-ai-2026.git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.github-code-nanochatbpe-1B
github-code-nanochatbpe-1B
GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
1,000,000,000
val.bin
val
10,000,000
train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.git-commits
Dataset: dataset.jsonl
Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label.
Features
Field
Type
Description
Stats
text
string
Commit message first line, conventional prefix stripped
—
files_count
int
Number of files changed
mean 4.3, median 1, max 300
additions
int
Lines added
mean 88, median 6, max 187K
deletions
int
Lines deleted
mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.cuda-nsys-training
Qwythos Nsight Systems Profiling Agent Dataset
Multi-turn GPU profiling agent trajectories for fine-tuning Qwythos-9B (and similar tool-calling models) on NVIDIA Nsight Systems (nsys) + CUDA-L1 / KernelBench workloads.
Generated autonomously on an RTX 5090 by the model itself driving real profiling tools for ~33 hours.
Code: ai-hpc/prof-dataset-gen
Stats
Split
Rows
Notes
train
5,884
Accepted episodes (quality ≥ 0.55)
eval
309
5% holdout from accepted… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/cuda-nsys-training.github-issuesannotations_creators:
other
language_creators:
crowdsourced
languages:
en-US
licenses:
other-my-license
multilinguality:
monolingual
pretty_name: HuggingFace Github Issues
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
github-trending
GitHub Trending: топ репозиториев за неделю (2026-09-28)
Открытые данные GitHub API: топ-10 репозиториев за 7 дней.
ai-github-issues-2026
ai-github-issues-2026
Real GitHub issues from AI repos
Records: 72 | Real data | Updated: 2026-10-06
API: GET https://api.legion-api.com/github-issues
Bundle: gemmo.gumroad.com/l/mdevxu
Gated — auto-approved.
github-repo-scraper
GitHub Repo Scraper · Repositories, Stars, Topics & Languages
Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication.
Rows in this dataset
2,481
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.ai-github-discussions-2026
ai-github-discussions-2026
GitHub discussions. GraphQL
Records: 15 | Real data | Updated: 2026-10-06
API: GET https://api.legion-api.com/github-discussions
Bundle: gemmo.gumroad.com/l/mdevxu
Gated — auto-approved.
ai-github-trending-2026
ai-github-trending-2026
Real GitHub AI repos. source: api.github.com
Records: 169 | 100% real data | Updated: 2026-10-03
API: GET https://api.legion-api.com/github-trending
Bundle: gemmo.gumroad.com/l/mdevxu
Access gated — auto-approved.
transformers-github-issuesbhagavad-gita-with_life_lesson
bhagavad-gita-lifelesson Dataset
A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format:
First: Sanskrit chanting
Then: Hindi meaning (हिन्दी अर्थ)
Then: Life lesson (जीवन-पाठ)
(Transliteration and English translation have been removed).
🎧 Example Representation (Verse 2.47)
🎧 Verse 2.47
First: Sanskrit chanting
कर्मण्येवाधिकारस्ते मा फलेषु कदाचन
मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.ai-github-prs-2026
ai-github-prs-2026
Real merged GitHub PRs — what changed in AI codebases this week
Records: 28 | 100% real data | Updated: 2026-10-03
API: GET https://api.legion-api.com/github-prs
Bundle: gemmo.gumroad.com/l/mdevxu
Access gated — auto-approved.
github-issuesannotations_creators:
machine-generated
language_creators:
machine-generated
languages:
english
licenses:
unknown
multilinguality:
monolingual
pretty_name: This dataset contains issues from Hugging Face Github
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
github-issuesCode-Githubbhagavad-gita-verses-sanskrit-translations
Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset
A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes:
📜 Original Sanskrit slokas
🔡 IAST transliteration
🌐 Multiple English & Hindi translations
🧠 Traditional commentaries from many teachers
🔢 Structured metadata (chapter, verse, IDs, authors)
This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.github-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/planhanasan/github-issues.github-issuesgithub-issuesgithub-issuesannotations_creators:
no-annotation
language_creators:
found
languages:
en
licenses:
unknown
multilinguality:
monolingual
pretty_name: Practice
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
Dataset Card for [Needs More Information]
Dataset Summary
For Practice
Supported Tasks and Leaderboards
Classification… See the full description on the dataset page: https://huggingface.co/datasets/peterhsu/github-issues.
