blindsubmissions/GH_text2code
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming language pairs. Namely, text is paired with code snippets for: Python, Java, JavaScript, and Go. The data is curated via an automated filtering pipeline from source files within The Stack. Supported Tasks This dataset can be used to finetune models for code-to-text and/or text-to-code models, both on information retrieval or… See the full description on the dataset page: https://huggingface.co/datasets/blindsubmissions/GH_text2code.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face