semeru/code-text-ruby
Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/ruby in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-ruby.
551
1---2license: mit3Programminglanguage: "ruby"4version: "N/A"5Date: "Codesearchnet(Jun 2020 - paper release date)"6Contaminated: "Very Likely"7Size: "Standar Tokenizer (TreeSitter)"8 9---10 11### Dataset is imported from CodeXGLUE and pre-processed using their script.12 13# Where to find in Semeru:14The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/ruby in Semeru15 16 17# CodeXGLUE -- Code-To-Text18 19## Task Definition20 21The task is to generate natural language comments for a code, and evaluted by [smoothed bleu-4](https://www.aclweb.org/anthology/C04-1072.pdf) score.22 23## Dataset24 25The dataset we use comes from [CodeSearchNet](https://arxiv.org/pdf/1909.09436.pdf) and we filter the dataset as the following:26 27- Remove examples that codes cannot be parsed into an abstract syntax tree.28- Remove examples that #tokens of documents is < 3 or >25629- Remove examples that documents contain special tokens (e.g. <img ...> or https:...)30- Remove examples that documents are not English.31 32 33### Data Format34 35After preprocessing dataset, you can obtain three .jsonl files, i.e. train.jsonl, valid.jsonl, test.jsonl36 37For each file, each line in the uncompressed file represents one function. One row is illustrated below.38 39 - **repo:** the owner/repo40 41 - **path:** the full path to the original file42 43 - **func_name:** the function or method name44 45 - **original_string:** the raw string before tokenization or parsing46 47 - **language:** the programming language48 49 - **code/function:** the part of the `original_string` that is code50 51 - **code_tokens/function_tokens:** tokenized version of `code`52 53 - **docstring:** the top-level comment or docstring, if it exists in the original string54 55 - **docstring_tokens:** tokenized version of `docstring`56 57### Data Statistic58 59| Programming Language | Training | Dev | Test |60| :------------------- | :------: | :----: | :----: |61| Ruby | 24,927 | 1,400 | 1,261 |62 63## Reference64<pre><code>@article{husain2019codesearchnet,65 title={Codesearchnet challenge: Evaluating the state of semantic code search},66 author={Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},67 journal={arXiv preprint arXiv:1909.09436},68 year={2019}69}</code></pre>70 