datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_docstring_ast_corpus
Overview
This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their
publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks.
Sources
The dataset was gathered from various GitHub repos sampled from this repo by Vinta.
The 26 repos are:
matplotlib
pytorch
cryptography
django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.powershell-code-corpus
PowerShell Code Corpus
A dataset of PowerShell scripts collected from public GitHub repositories.
Content
4,632 files from popular PowerShell repositories on GitHub
Sourced from repos with the highest star counts (quality signal)
Each record contains the raw script text plus metadata
Fields
Field
Description
source
Always github
repo
owner/repo slug
repo_url
Full GitHub URL
path
File path within the repo
language
Always PowerShell
license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.code-cpt-corpus
Dataset Card for code-cpt-corpus
Dataset Summary
ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/train.jsonl
Licensing Information
ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.code-quality-corpus
CatQualia code-quality corpus — semantic smell classes with before/after fixes
39,383 rows · 25,541,792 bytes · JSON Lines, one object per line.
What this is
Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.ruby-code-corpus
Ruby Code Corpus
A large corpus of Ruby source code collected from public GitHub repositories.
Content
294,074 files from 3,920+ Ruby repositories on GitHub
Sourced from repos ranked by star count (quality signal)
Filtered: removed files under 200 bytes (trivial/empty files)
Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby
Fields
Field
Description
source
Always github
repo
owner/repo slug
repo_url
Full GitHub URL
path
File path… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-corpus.
