Team Ai
Datasetpublic

codeparrot/github-jupyter-text-code-pairs

This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
7likes122downloads
README.md20 linesDownload Raw Back to root
1---2annotations_creators: []3language:4- code5license:6- other7multilinguality:8- monolingual9size_categories:10- unknown11task_categories:12- text-generation13task_ids:14- language-modeling15pretty_name: github-jupyter-text-code-pairs16---17 18This is a parsed version of [github-jupyter-parsed](https://huggingface.co/datasets/codeparrot/github-jupyter-parsed), with markdown and code pairs. We provide the preprocessing script in [preprocessing.py](https://huggingface.co/datasets/codeparrot/github-jupyter-parsed-v2/blob/main/preprocessing.py). The data is deduplicated and consists of 451662 examples. 19 20For similar datasets with text and Python code, there is [CoNaLa](https://huggingface.co/datasets/neulab/conala) benchmark from StackOverflow, with some samples curated by annotators.