Team Ai
Datasetpublic

blindsubmissions/GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming language pairs. Namely, text is paired with code snippets for: Python, Java, JavaScript, and Go. The data is curated via an automated filtering pipeline from source files within The Stack. Supported Tasks This dataset can be used to finetune models for code-to-text and/or text-to-code models, both on information retrieval or… See the full description on the dataset page: https://huggingface.co/datasets/blindsubmissions/GH_text2code.

sourceHugging Faceupdated 3y agoView on Hugging Face
4likes548downloads
README.md134 linesDownload Raw Back to root
1---2dataset_info:3  features:4  - name: identifier5    dtype: string6  - name: parameters7    dtype: string8  - name: docstring9    dtype: string10  - name: docstring_summary11    dtype: string12  - name: function13    dtype: string14  - name: function_tokens15    sequence: string16  - name: start_point17    sequence: int6418  - name: end_point19    sequence: int6420  - name: language21    dtype: string22  - name: docstring_language23    dtype: string24  - name: docstring_language_predictions25    dtype: string26  - name: is_langid_reliable27    dtype: string28  splits:29  - name: python_gh30    num_bytes: 3630076042331    num_examples: 1500000232  - name: java_gh33    num_bytes: 2161305711034    num_examples: 1500001435  - name: go_gh36    num_bytes: 2255974193737    num_examples: 1500007838  - name: javascript_gh39    num_bytes: 389568831140    num_examples: 200004041  download_size: 16632449942  dataset_size: 8436924778143task_categories:44- translation45- summarization46- text2text-generation47language:48- en49tags:50- code51size_categories:52- 10M<n<100M53---54# Docstring to code data55 56## Dataset Summary57This dataset contains pairs of English text and code from multiple programming language pairs. Namely, text is paired with code snippets for: Python, Java, JavaScript, and Go. The data is curated via an automated filtering pipeline from source files within [The Stack](https://huggingface.co/datasets/bigcode/the-stack).58 59## Supported Tasks60This dataset can be used to finetune models  for code-to-text and/or text-to-code models, both on information retrieval or conditional generation settings.61 62## Splits63 64```python65DATA_SPLITS = {"python_gh", "java_gh", "javascript_gh", "go_gh"}66```67 68## How to get the data with a given programming language69 70```python71from datasets import load_dataset72 73def get_dataset(prog_lang):74 75  test_data = load_dataset("blindsubmissions/GH_text2code", split=prog_lang)76 77  return test_data78```79 80## Dataset Structure81 82### Data Instances83Each data instance corresponds to function/methods occurring in licensed files that compose The Stack. That is, files with permissive licences collected from GitHub.84 85### Relevant Data Fields86 87- identifier (string): Function/method name.88- parameters (string): Function parameters.89- return_statement (string): Return statement if found during parsing.90- docstring (string): Complete docstring content.91- docstring_summary (string): Summary/processed docstring dropping args and return statements.92- function (string): Actual function/method content.93- argument_list (null): List of arguments.94- language (string): Programming language of the function.95- type (string): Return type if found during parsing.96 97## Summary of data curation pipeline98 99- Filtering out repositories that appear in [CodeSearchNet](https://huggingface.co/datasets/code_search_net).100- Filtering the files that belong to the programming languages of interest.101- Pre-filtering the files that likely contain text in the natural languages of interest.102- AST parsing with [Tree-sitter](\url{https://tree-sitter.github.io/tree-sitter/).103- Perform language identification of docstrings in the resulting set of functions/methods and select the ones classified as English via majority voting.104 105## Social Impact of the dataset106 107This dataset is released with the aim to increase the availability of training data available to the NLP for code research community by providing text/code paired data. We expect this data to help enable more accurate information retrieval systems and text-to-code or code-to-text summarization.108 109As a subset of The Stack, this dataset inherits de-risking efforts carried out when that dataset was built, though we highlight risks exist and malicious use of the data could exist such as, for instance, to aid on creation of malicious code. We highlight however that this is a risk shared by any code dataset made openly available.110 111Moreover, we remark that the data may contain harmful or offensive language, which could be learned by models trained on it.112 113## Discussion of Biases114 115The data is collected from GitHub and naturally occurring text on that platform. As a consequence, certain languages are more or less likely to contain well documented code and, as such, resulting data will not be uniformly represented in terms of their programing languages.116 117## Known limitations118 119The dataset can be expanded to further improve its coverage.120Moreover, we use text naturally occurring as comments or docstrings as opposed to human annotators. As such, resulting data will have high variance in terms of quality depending on practices of sub-communities of software developers. However, we remark that the task our evaluation dataset defines is reflective of what searching on a real codebase would look like.121Finally, we note that some imbalance on data is observed due to the same reason since certain languages are more or less likely to contain well documented code.122 123## Maintenance plan:124 125The data will be kept up to date by following The Stack releases. We should rerun our pipeline for every new release and add non-overlapping new content to both training and testing partitions of our data. 126 127This is so that we carry over opt-out updates and include fresh repos.128 129## Update plan:130 131  - Cover all 6 programming languages from CodeSearchNet.132  133## Licensing Information134M2CRB is a subset filtered and pre-processed from [The Stack](https://huggingface.co/datasets/bigcode/the-stack), a collection of source code from repositories with various licenses. Any use of all or part of the code gathered in M2CRB must abide by the terms of the original licenses.