Team Ai
Datasetpublic

semeru/code-code-CodeRefinement-Java-Small

Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-refinement/data/small in Semeru Task Definition Code refinement aims to automatically fix bugs in the code, which can contribute to reducing the cost of bug-fixes for developers. In CodeXGLUE, given a piece of Java code with bugs, the task is to remove the bugs to… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeRefinement-Java-Small.

sourceHugging Facemitupdated 4y agoView on Hugging Face
0likes58downloads
README.md52 linesDownload Raw Back to root
1---2license: mit3Programminglanguage: "Java"4version: "N/A"5Date: "May 2019 paper release date for https://arxiv.org/pdf/1812.08693.pdf"6Contaminated: "Very Likely"7Size: "Standard Tokenizer"8 9---10 11 12 13### Dataset is imported from CodeXGLUE and pre-processed using their script.14 15# Where to find in Semeru:16The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-refinement/data/small in Semeru17 18## Task Definition19 20Code refinement aims to automatically fix bugs in the code, which can contribute to reducing the cost of bug-fixes for developers.21In CodeXGLUE, given a piece of Java code with bugs, the task is to remove the bugs to output the refined code. 22Models are evaluated by BLEU scores, accuracy (exactly match) and [CodeBLEU](https://github.com/microsoft/CodeXGLUE/blob/main/code-to-code-trans/CodeBLEU.MD).23 24## Dataset25 26We use the dataset released by this paper(https://arxiv.org/pdf/1812.08693.pdf). The source side is a Java function with bugs and the target side is the refined one. 27All the function and variable names are normalized. Their dataset contains two subsets ( i.e.small and medium) based on the function length. This dataset is small.28 29 30 31### Data Statistics32 33Data statistics of this dataset are shown in the below table:34 35|         | #Examples |36| ------- | :-------: | 37|         |   Small   |  38|  Train  |   46,680  |  39|  Valid  |    5,835  |   40|   Test  |    5,835  |   41 42# Reference43<pre><code>@article{tufano2019empirical,44  title={An empirical study on learning bug-fixing patches in the wild via neural machine translation},45  author={Tufano, Michele and Watson, Cody and Bavota, Gabriele and Penta, Massimiliano Di and White, Martin and Poshyvanyk, Denys},46  journal={ACM Transactions on Software Engineering and Methodology (TOSEM)},47  volume={28},48  number={4},49  pages={1--29},50  year={2019},51  publisher={ACM New York, NY, USA}52}</code></pre>