Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mir-2002 /python_code_docstring_ast_corpus Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.textsummarization10K<n<100K1 likes52 downloads1y agoHugging Face02NickIBrody /powershell-code-corpus PowerShell Code Corpus A dataset of PowerShell scripts collected from public GitHub repositories. Content 4,632 files from popular PowerShell repositories on GitHub Sourced from repos with the highest star counts (quality signal) Each record contains the raw script text plus metadata Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path File path within the repo language Always PowerShell license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.tabular1K<n<10K0 likes50 downloads6mo agoHugging Face03kkomyoeminaung /code-cpt-corpus Dataset Card for code-cpt-corpus Dataset Summary ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/train.jsonl Licensing Information ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.text100K<n<1M0 likes49 downloads3mo agoHugging Face04CatQualia /code-quality-corpusgated CatQualia code-quality corpus — semantic smell classes with before/after fixes 39,383 rows · 25,541,792 bytes · JSON Lines, one object per line. What this is Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.tabular10K<n<100K0 likes31 downloads26d agoHugging Face05NickIBrody /ruby-code-corpus Ruby Code Corpus A large corpus of Ruby source code collected from public GitHub repositories. Content 294,074 files from 3,920+ Ruby repositories on GitHub Sourced from repos ranked by star count (quality signal) Filtered: removed files under 200 bytes (trivial/empty files) Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path File path… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-corpus.tabular100K<n<1M0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.