Team Ai
Datasetpublic

codeparrot/github-jupyter-text-code-pairs

This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
7likes122downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
codeparrot/github-jupyter-text-code-pairs · Team Ai