Team Ai
Datasetpublic

codeparrot/github-jupyter-code-to-text

Dataset description This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs from this dataset that were originally code and markdown cells in Jupyter Notebooks. The content of each example the following: [CODE] """ Explanation: [TEXT] End of explanation """ [CODE] """ Explanation: [TEXT] End of explanation """ ... How to use it from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
27likes228downloads
README.md42 linesDownload Raw Back to root
1---2license: apache-2.03task_categories:4- text-generation5tags:6- code7size_categories:8- 10K<n<100K9language:10- en11---12 13# Dataset description14This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs 15from this [dataset](https://huggingface.co/datasets/codeparrot/github-jupyter-text-code-pairs) that were originally code and markdown cells in Jupyter Notebooks.16 17The content of each example the following:18````19[CODE]20"""21Explanation: [TEXT]22End of explanation23"""24[CODE]25"""26Explanation: [TEXT]27End of explanation28"""29...30````31# How to use it32```python33from datasets import load_dataset34 35ds = load_dataset("codeparrot/github-jupyter-code-to-text", split="train")36````37````38Dataset({39    features: ['repo_name', 'path', 'license', 'content'],40    num_rows: 4745241})42````