semeru/Text-Code-concode-Java
Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/text-to-code/concode in Semeru CodeXGLUE -- Text2Code Generation Here are the dataset and pipeline for text-to-code generation task. Task Definition Generate source code of class member functions in Java, given natural language description and class environment. Class… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Text-Code-concode-Java.
5127
1---2license: mit3Programminglanguage: "Java"4version: "N/A"5Date: "2018 paper https://aclanthology.org/D18-1192.pdf"6Contaminated: "Very Likely"7Size: "Standard Tokenizer"8 9---10## Dataset is imported from CodeXGLUE and pre-processed using their script.11 12# Where to find in Semeru:13The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/text-to-code/concode in Semeru14 15 16# CodeXGLUE -- Text2Code Generation17 18Here are the dataset and pipeline for text-to-code generation task.19 20## Task Definition21 22Generate source code of class member functions in Java, given natural language description and class environment. Class environment is the programmatic context provided by the rest of the class, including other member variables and member functions in class. Models are evaluated by exact match and BLEU.23 24It's a challenging task because the desired code can vary greatly depending on the functionality the class provides. Models must (a) have a deep understanding of NL description and map the NL to environment variables, library API calls and user-defined methods in the class, and (b) decide on the structure of the resulting code.25 26 27## Dataset28 29### Concode dataset30We use concode dataset which is a widely used code generation dataset from Iyer's EMNLP 2018 paper [Mapping Language to Code in Programmatic Context](https://www.aclweb.org/anthology/D18-1192.pdf).31 32We have downloaded his published dataset and followed his preprocessed script. You can find the preprocessed data in `dataset/concode` directory.33 34Data statistics of concode dataset are shown in the below table:35 36| | #Examples |37| ------- | :---------: |38| Train | 100,000 |39| Dev | 2,000 |40| Test | 2,000 |41 42### Data Format43 44Code corpus are saved in json lines format files. one line is a json object:45```46{47 "nl": "Increment this vector in this place. con_elem_sep double[] vecElement con_elem_sep double[] weights con_func_sep void add(double)",48 "code": "public void inc ( ) { this . add ( 1 ) ; }"49}50```51 52`nl` combines natural language description and class environment. Elements in class environment are seperated by special tokens like `con_elem_sep` and `con_func_sep`.53 54 55## Reference56 57 58<pre><code>@article{iyer2018mapping,59 title={Mapping language to code in programmatic context},60 author={Iyer, Srinivasan and Konstas, Ioannis and Cheung, Alvin and Zettlemoyer, Luke},61 journal={arXiv preprint arXiv:1808.09588},62 year={2018}63}</code></pre>64 