Team Ai
Datasetpublic

google/code_x_glue_cc_code_to_code_trans

Dataset Card for "code_x_glue_cc_code_to_code_trans" Dataset Summary CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/). We collect both the Java and C# versions of the codes and find the… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_to_code_trans.

sourceHugging Facec-udaupdated 3y agoView on Hugging Face
17likes625downloads
README.md221 linesDownload Raw Back to root
1---2annotations_creators:3- expert-generated4language_creators:5- found6language:7- code8license:9- c-uda10multilinguality:11- other-programming-languages12size_categories:13- 10K<n<100K14source_datasets:15- original16task_categories:17- translation18task_ids: []19pretty_name: CodeXGlueCcCodeToCodeTrans20tags:21- code-to-code22dataset_info:23  features:24  - name: id25    dtype: int3226  - name: java27    dtype: string28  - name: cs29    dtype: string30  splits:31  - name: train32    num_bytes: 437264133    num_examples: 1030034  - name: validation35    num_bytes: 22640736    num_examples: 50037  - name: test38    num_bytes: 41858739    num_examples: 100040  download_size: 206476441  dataset_size: 501763542configs:43- config_name: default44  data_files:45  - split: train46    path: data/train-*47  - split: validation48    path: data/validation-*49  - split: test50    path: data/test-*51---52# Dataset Card for "code_x_glue_cc_code_to_code_trans"53 54## Table of Contents55- [Dataset Description](#dataset-description)56  - [Dataset Summary](#dataset-summary)57  - [Supported Tasks and Leaderboards](#supported-tasks)58  - [Languages](#languages)59- [Dataset Structure](#dataset-structure)60  - [Data Instances](#data-instances)61  - [Data Fields](#data-fields)62  - [Data Splits](#data-splits-sample-size)63- [Dataset Creation](#dataset-creation)64  - [Curation Rationale](#curation-rationale)65  - [Source Data](#source-data)66  - [Annotations](#annotations)67  - [Personal and Sensitive Information](#personal-and-sensitive-information)68- [Considerations for Using the Data](#considerations-for-using-the-data)69  - [Social Impact of Dataset](#social-impact-of-dataset)70  - [Discussion of Biases](#discussion-of-biases)71  - [Other Known Limitations](#other-known-limitations)72- [Additional Information](#additional-information)73  - [Dataset Curators](#dataset-curators)74  - [Licensing Information](#licensing-information)75  - [Citation Information](#citation-information)76  - [Contributions](#contributions)77 78## Dataset Description79 80- **Homepage:** https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans81- **Paper:** https://arxiv.org/abs/2102.0466482 83### Dataset Summary84 85CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans86 87The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/).88 89We collect both the Java and C# versions of the codes and find the parallel functions. After removing duplicates and functions with the empty body, we split the whole dataset into training, validation and test sets.90 91### Supported Tasks and Leaderboards92 93- `machine-translation`: The dataset can be used to train a model for translating code in Java to C# and vice versa.94 95### Languages96 97- Java **programming** language98- C# **programming** language99 100## Dataset Structure101 102### Data Instances103 104An example of 'validation' looks as follows.105```106{107    "cs": "public DVRecord(RecordInputStream in1){_option_flags = in1.ReadInt();_promptTitle = ReadUnicodeString(in1);_errorTitle = ReadUnicodeString(in1);_promptText = ReadUnicodeString(in1);_errorText = ReadUnicodeString(in1);int field_size_first_formula = in1.ReadUShort();_not_used_1 = in1.ReadShort();_formula1 = NPOI.SS.Formula.Formula.Read(field_size_first_formula, in1);int field_size_sec_formula = in1.ReadUShort();_not_used_2 = in1.ReadShort();_formula2 = NPOI.SS.Formula.Formula.Read(field_size_sec_formula, in1);_regions = new CellRangeAddressList(in1);}\n", 108    "id": 0, 109    "java": "public DVRecord(RecordInputStream in) {_option_flags = in.readInt();_promptTitle = readUnicodeString(in);_errorTitle = readUnicodeString(in);_promptText = readUnicodeString(in);_errorText = readUnicodeString(in);int field_size_first_formula = in.readUShort();_not_used_1 = in.readShort();_formula1 = Formula.read(field_size_first_formula, in);int field_size_sec_formula = in.readUShort();_not_used_2 = in.readShort();_formula2 = Formula.read(field_size_sec_formula, in);_regions = new CellRangeAddressList(in);}\n"110}111```112 113### Data Fields114 115In the following each data field in go is explained for each config. The data fields are the same among all splits.116 117#### default118 119|field name| type |         description         |120|----------|------|-----------------------------|121|id        |int32 | Index of the sample         |122|java      |string| The java version of the code|123|cs        |string| The C# version of the code  |124 125### Data Splits126 127| name  |train|validation|test|128|-------|----:|---------:|---:|129|default|10300|       500|1000|130 131## Dataset Creation132 133### Curation Rationale134 135[More Information Needed]136 137### Source Data138 139#### Initial Data Collection and Normalization140 141[More Information Needed]142 143#### Who are the source language producers?144 145[More Information Needed]146 147### Annotations148 149#### Annotation process150 151[More Information Needed]152 153#### Who are the annotators?154 155[More Information Needed]156 157### Personal and Sensitive Information158 159[More Information Needed]160 161## Considerations for Using the Data162 163### Social Impact of Dataset164 165[More Information Needed]166 167### Discussion of Biases168 169[More Information Needed]170 171### Other Known Limitations172 173[More Information Needed]174 175## Additional Information176 177### Dataset Curators178 179https://github.com/microsoft, https://github.com/madlag180 181### Licensing Information182 183Computational Use of Data Agreement (C-UDA) License.184 185### Citation Information186 187```188@article{DBLP:journals/corr/abs-2102-04664,189  author    = {Shuai Lu and190               Daya Guo and191               Shuo Ren and192               Junjie Huang and193               Alexey Svyatkovskiy and194               Ambrosio Blanco and195               Colin B. Clement and196               Dawn Drain and197               Daxin Jiang and198               Duyu Tang and199               Ge Li and200               Lidong Zhou and201               Linjun Shou and202               Long Zhou and203               Michele Tufano and204               Ming Gong and205               Ming Zhou and206               Nan Duan and207               Neel Sundaresan and208               Shao Kun Deng and209               Shengyu Fu and210               Shujie Liu},211  title     = {CodeXGLUE: {A} Machine Learning Benchmark Dataset for Code Understanding212               and Generation},213  journal   = {CoRR},214  volume    = {abs/2102.04664},215  year      = {2021}216}217```218 219### Contributions220 221Thanks to @madlag (and partly also @ncoop57) for adding this dataset.