Team Ai
Datasetpublic

blindsubmissions/GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming language pairs. Namely, text is paired with code snippets for: Python, Java, JavaScript, and Go. The data is curated via an automated filtering pipeline from source files within The Stack. Supported Tasks This dataset can be used to finetune models for code-to-text and/or text-to-code models, both on information retrieval or… See the full description on the dataset page: https://huggingface.co/datasets/blindsubmissions/GH_text2code.

sourceHugging Faceupdated 3y agoView on Hugging Face
4likes548downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
blindsubmissions/GH_text2code · Team Ai