google/code_x_glue_tc_text_to_code
Dataset Card for "code_x_glue_tc_text_to_code" Dataset Summary CodeXGLUE text-to-code dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/. Supported Tasks and Leaderboards machine-translation: The dataset can be used to train a model for generating Java code from an English… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_text_to_code.
30796
1---2annotations_creators:3- found4language_creators:5- found6language:7- code8- en9license:10- c-uda11multilinguality:12- other-programming-languages13size_categories:14- 100K<n<1M15source_datasets:16- original17task_categories:18- translation19task_ids: []20pretty_name: CodeXGlueTcTextToCode21tags:22- text-to-code23dataset_info:24 features:25 - name: id26 dtype: int3227 - name: nl28 dtype: string29 - name: code30 dtype: string31 splits:32 - name: train33 num_bytes: 9622553134 num_examples: 10000035 - name: validation36 num_bytes: 174974337 num_examples: 200038 - name: test39 num_bytes: 160929840 num_examples: 200041 download_size: 3425835442 dataset_size: 9958457243configs:44- config_name: default45 data_files:46 - split: train47 path: data/train-*48 - split: validation49 path: data/validation-*50 - split: test51 path: data/test-*52---53# Dataset Card for "code_x_glue_tc_text_to_code"54 55## Table of Contents56- [Dataset Description](#dataset-description)57 - [Dataset Summary](#dataset-summary)58 - [Supported Tasks and Leaderboards](#supported-tasks)59 - [Languages](#languages)60- [Dataset Structure](#dataset-structure)61 - [Data Instances](#data-instances)62 - [Data Fields](#data-fields)63 - [Data Splits](#data-splits-sample-size)64- [Dataset Creation](#dataset-creation)65 - [Curation Rationale](#curation-rationale)66 - [Source Data](#source-data)67 - [Annotations](#annotations)68 - [Personal and Sensitive Information](#personal-and-sensitive-information)69- [Considerations for Using the Data](#considerations-for-using-the-data)70 - [Social Impact of Dataset](#social-impact-of-dataset)71 - [Discussion of Biases](#discussion-of-biases)72 - [Other Known Limitations](#other-known-limitations)73- [Additional Information](#additional-information)74 - [Dataset Curators](#dataset-curators)75 - [Licensing Information](#licensing-information)76 - [Citation Information](#citation-information)77 - [Contributions](#contributions)78 79## Dataset Description80 81- **Homepage:** https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code82 83### Dataset Summary84 85CodeXGLUE text-to-code dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code86 87The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/.88 89### Supported Tasks and Leaderboards90 91- `machine-translation`: The dataset can be used to train a model for generating Java code from an **English** natural language description.92 93### Languages94 95- Java **programming** language96 97## Dataset Structure98 99### Data Instances100 101An example of 'train' looks as follows.102```103{104 "code": "boolean function ( ) { return isParsed ; }", 105 "id": 0, 106 "nl": "check if details are parsed . concode_field_sep Container parent concode_elem_sep boolean isParsed concode_elem_sep long offset concode_elem_sep long contentStartPosition concode_elem_sep ByteBuffer deadBytes concode_elem_sep boolean isRead concode_elem_sep long memMapSize concode_elem_sep Logger LOG concode_elem_sep byte[] userType concode_elem_sep String type concode_elem_sep ByteBuffer content concode_elem_sep FileChannel fileChannel concode_field_sep Container getParent concode_elem_sep byte[] getUserType concode_elem_sep void readContent concode_elem_sep long getOffset concode_elem_sep long getContentSize concode_elem_sep void getContent concode_elem_sep void setDeadBytes concode_elem_sep void parse concode_elem_sep void getHeader concode_elem_sep long getSize concode_elem_sep void parseDetails concode_elem_sep String getType concode_elem_sep void _parseDetails concode_elem_sep String getPath concode_elem_sep boolean verify concode_elem_sep void setParent concode_elem_sep void getBox concode_elem_sep boolean isSmallBox"107}108```109 110### Data Fields111 112In the following each data field in go is explained for each config. The data fields are the same among all splits.113 114#### default115 116|field name| type | description |117|----------|------|---------------------------------------------|118|id |int32 | Index of the sample |119|nl |string| The natural language description of the task|120|code |string| The programming source code for the task |121 122### Data Splits123 124| name |train |validation|test|125|-------|-----:|---------:|---:|126|default|100000| 2000|2000|127 128## Dataset Creation129 130### Curation Rationale131 132[More Information Needed]133 134### Source Data135 136#### Initial Data Collection and Normalization137 138[More Information Needed]139 140#### Who are the source language producers?141 142[More Information Needed]143 144### Annotations145 146#### Annotation process147 148[More Information Needed]149 150#### Who are the annotators?151 152[More Information Needed]153 154### Personal and Sensitive Information155 156[More Information Needed]157 158## Considerations for Using the Data159 160### Social Impact of Dataset161 162[More Information Needed]163 164### Discussion of Biases165 166[More Information Needed]167 168### Other Known Limitations169 170[More Information Needed]171 172## Additional Information173 174### Dataset Curators175 176https://github.com/microsoft, https://github.com/madlag177 178### Licensing Information179 180Computational Use of Data Agreement (C-UDA) License.181 182### Citation Information183 184```185@article{iyer2018mapping,186 title={Mapping language to code in programmatic context},187 author={Iyer, Srinivasan and Konstas, Ioannis and Cheung, Alvin and Zettlemoyer, Luke},188 journal={arXiv preprint arXiv:1808.09588},189 year={2018}190}191```192 193### Contributions194 195Thanks to @madlag (and partly also @ncoop57) for adding this dataset.