codeparrot/github-jupyter-code-to-text
Dataset description This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs from this dataset that were originally code and markdown cells in Jupyter Notebooks. The content of each example the following: [CODE] """ Explanation: [TEXT] End of explanation """ [CODE] """ Explanation: [TEXT] End of explanation """ ... How to use it from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.
27228
1---2license: apache-2.03task_categories:4- text-generation5tags:6- code7size_categories:8- 10K<n<100K9language:10- en11---12 13# Dataset description14This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs 15from this [dataset](https://huggingface.co/datasets/codeparrot/github-jupyter-text-code-pairs) that were originally code and markdown cells in Jupyter Notebooks.16 17The content of each example the following:18````19[CODE]20"""21Explanation: [TEXT]22End of explanation23"""24[CODE]25"""26Explanation: [TEXT]27End of explanation28"""29...30````31# How to use it32```python33from datasets import load_dataset34 35ds = load_dataset("codeparrot/github-jupyter-code-to-text", split="train")36````37````38Dataset({39 features: ['repo_name', 'path', 'license', 'content'],40 num_rows: 4745241})42````