Team Ai
Datasetpublic

Mir-2002/python_code_docstring_ast_corpus

Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes63downloads
settings

This repository belongs to Mir-2002 on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namepython_code_docstring_ast_corpus
visibilitypublic
licencenot set
gatedno
ownerMir-2002
Account settings
Mir-2002/python_code_docstring_ast_corpus · Team Ai