google/code_x_glue_cc_code_to_code_trans
Dataset Card for "code_x_glue_cc_code_to_code_trans" Dataset Summary CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/). We collect both the Java and C# versions of the codes and find the… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_to_code_trans.
17625
1---2annotations_creators:3- expert-generated4language_creators:5- found6language:7- code8license:9- c-uda10multilinguality:11- other-programming-languages12size_categories:13- 10K<n<100K14source_datasets:15- original16task_categories:17- translation18task_ids: []19pretty_name: CodeXGlueCcCodeToCodeTrans20tags:21- code-to-code22dataset_info:23 features:24 - name: id25 dtype: int3226 - name: java27 dtype: string28 - name: cs29 dtype: string30 splits:31 - name: train32 num_bytes: 437264133 num_examples: 1030034 - name: validation35 num_bytes: 22640736 num_examples: 50037 - name: test38 num_bytes: 41858739 num_examples: 100040 download_size: 206476441 dataset_size: 501763542configs:43- config_name: default44 data_files:45 - split: train46 path: data/train-*47 - split: validation48 path: data/validation-*49 - split: test50 path: data/test-*51---52# Dataset Card for "code_x_glue_cc_code_to_code_trans"53 54## Table of Contents55- [Dataset Description](#dataset-description)56 - [Dataset Summary](#dataset-summary)57 - [Supported Tasks and Leaderboards](#supported-tasks)58 - [Languages](#languages)59- [Dataset Structure](#dataset-structure)60 - [Data Instances](#data-instances)61 - [Data Fields](#data-fields)62 - [Data Splits](#data-splits-sample-size)63- [Dataset Creation](#dataset-creation)64 - [Curation Rationale](#curation-rationale)65 - [Source Data](#source-data)66 - [Annotations](#annotations)67 - [Personal and Sensitive Information](#personal-and-sensitive-information)68- [Considerations for Using the Data](#considerations-for-using-the-data)69 - [Social Impact of Dataset](#social-impact-of-dataset)70 - [Discussion of Biases](#discussion-of-biases)71 - [Other Known Limitations](#other-known-limitations)72- [Additional Information](#additional-information)73 - [Dataset Curators](#dataset-curators)74 - [Licensing Information](#licensing-information)75 - [Citation Information](#citation-information)76 - [Contributions](#contributions)77 78## Dataset Description79 80- **Homepage:** https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans81- **Paper:** https://arxiv.org/abs/2102.0466482 83### Dataset Summary84 85CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans86 87The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/).88 89We collect both the Java and C# versions of the codes and find the parallel functions. After removing duplicates and functions with the empty body, we split the whole dataset into training, validation and test sets.90 91### Supported Tasks and Leaderboards92 93- `machine-translation`: The dataset can be used to train a model for translating code in Java to C# and vice versa.94 95### Languages96 97- Java **programming** language98- C# **programming** language99 100## Dataset Structure101 102### Data Instances103 104An example of 'validation' looks as follows.105```106{107 "cs": "public DVRecord(RecordInputStream in1){_option_flags = in1.ReadInt();_promptTitle = ReadUnicodeString(in1);_errorTitle = ReadUnicodeString(in1);_promptText = ReadUnicodeString(in1);_errorText = ReadUnicodeString(in1);int field_size_first_formula = in1.ReadUShort();_not_used_1 = in1.ReadShort();_formula1 = NPOI.SS.Formula.Formula.Read(field_size_first_formula, in1);int field_size_sec_formula = in1.ReadUShort();_not_used_2 = in1.ReadShort();_formula2 = NPOI.SS.Formula.Formula.Read(field_size_sec_formula, in1);_regions = new CellRangeAddressList(in1);}\n", 108 "id": 0, 109 "java": "public DVRecord(RecordInputStream in) {_option_flags = in.readInt();_promptTitle = readUnicodeString(in);_errorTitle = readUnicodeString(in);_promptText = readUnicodeString(in);_errorText = readUnicodeString(in);int field_size_first_formula = in.readUShort();_not_used_1 = in.readShort();_formula1 = Formula.read(field_size_first_formula, in);int field_size_sec_formula = in.readUShort();_not_used_2 = in.readShort();_formula2 = Formula.read(field_size_sec_formula, in);_regions = new CellRangeAddressList(in);}\n"110}111```112 113### Data Fields114 115In the following each data field in go is explained for each config. The data fields are the same among all splits.116 117#### default118 119|field name| type | description |120|----------|------|-----------------------------|121|id |int32 | Index of the sample |122|java |string| The java version of the code|123|cs |string| The C# version of the code |124 125### Data Splits126 127| name |train|validation|test|128|-------|----:|---------:|---:|129|default|10300| 500|1000|130 131## Dataset Creation132 133### Curation Rationale134 135[More Information Needed]136 137### Source Data138 139#### Initial Data Collection and Normalization140 141[More Information Needed]142 143#### Who are the source language producers?144 145[More Information Needed]146 147### Annotations148 149#### Annotation process150 151[More Information Needed]152 153#### Who are the annotators?154 155[More Information Needed]156 157### Personal and Sensitive Information158 159[More Information Needed]160 161## Considerations for Using the Data162 163### Social Impact of Dataset164 165[More Information Needed]166 167### Discussion of Biases168 169[More Information Needed]170 171### Other Known Limitations172 173[More Information Needed]174 175## Additional Information176 177### Dataset Curators178 179https://github.com/microsoft, https://github.com/madlag180 181### Licensing Information182 183Computational Use of Data Agreement (C-UDA) License.184 185### Citation Information186 187```188@article{DBLP:journals/corr/abs-2102-04664,189 author = {Shuai Lu and190 Daya Guo and191 Shuo Ren and192 Junjie Huang and193 Alexey Svyatkovskiy and194 Ambrosio Blanco and195 Colin B. Clement and196 Dawn Drain and197 Daxin Jiang and198 Duyu Tang and199 Ge Li and200 Lidong Zhou and201 Linjun Shou and202 Long Zhou and203 Michele Tufano and204 Ming Gong and205 Ming Zhou and206 Nan Duan and207 Neel Sundaresan and208 Shao Kun Deng and209 Shengyu Fu and210 Shujie Liu},211 title = {CodeXGLUE: {A} Machine Learning Benchmark Dataset for Code Understanding212 and Generation},213 journal = {CoRR},214 volume = {abs/2102.04664},215 year = {2021}216}217```218 219### Contributions220 221Thanks to @madlag (and partly also @ncoop57) for adding this dataset.