Team Ai
Datasetpublic

semeru/code-code-CodeRefinement-Java-Medium

Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-refinement/data/medium in Semeru Task Definition Code refinement aims to automatically fix bugs in the code, which can contribute to reducing the cost of bug-fixes for developers. In CodeXGLUE, given a piece of Java code with bugs, the task is to remove the bugs to… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeRefinement-Java-Medium.

sourceHugging Facemitupdated 4y agoView on Hugging Face
0likes120downloads
README.md52 linesDownload Raw Back to root
1---2license: mit3Programminglanguage: "Java"4version: "N/A"5Date: "May 2019 paper release date for https://arxiv.org/pdf/1812.08693.pdf"6Contaminated: "Very Likely"7Size: "Standard Tokenizer "8---9 10### Dataset is imported from CodeXGLUE and pre-processed using their script.11 12# Where to find in Semeru:13The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-refinement/data/medium in Semeru14 15## Task Definition16 17Code refinement aims to automatically fix bugs in the code, which can contribute to reducing the cost of bug-fixes for developers.18In CodeXGLUE, given a piece of Java code with bugs, the task is to remove the bugs to output the refined code. 19Models are evaluated by BLEU scores, accuracy (exactly match) and [CodeBLEU](https://github.com/microsoft/CodeXGLUE/blob/main/code-to-code-trans/CodeBLEU.MD).20 21## Dataset22 23We use the dataset released by this paper(https://arxiv.org/pdf/1812.08693.pdf). The source side is a Java function with bugs and the target side is the refined one. 24All the function and variable names are normalized. Their dataset contains two subsets ( i.e.small and medium) based on the function length. This dataset is medium.25 26 27 28### Data Statistics29 30Data statistics of this dataset are shown in the below table:31 32|          | #Examples |33| -------  | :-------: |34|          |   Medium  |35|  Train   |   52,364  |36|  Valid   |    6,545  |37|   Test   |    6,545  |38 39# Reference40<pre><code>@article{tufano2019empirical,41  title={An empirical study on learning bug-fixing patches in the wild via neural machine translation},42  author={Tufano, Michele and Watson, Cody and Bavota, Gabriele and Penta, Massimiliano Di and White, Martin and Poshyvanyk, Denys},43  journal={ACM Transactions on Software Engineering and Methodology (TOSEM)},44  volume={28},45  number={4},46  pages={1--29},47  year={2019},48  publisher={ACM New York, NY, USA}49}</code></pre>50 51 52