Team Ai
Datasetpublic

google/code_x_glue_cc_code_completion_line

Dataset Card for "code_x_glue_cc_code_completion_line" Dataset Summary CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity. We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_line.

sourceHugging Facec-udaupdated 3y agoView on Hugging Face
10likes457downloads
README.md238 linesDownload Raw Back to root
1---2annotations_creators:3- found4language_creators:5- found6language:7- code8license:9- c-uda10multilinguality:11- monolingual12size_categories:13- 1K<n<10K14- n<1K15source_datasets:16- original17task_categories:18- text-generation19- fill-mask20task_ids:21- slot-filling22pretty_name: CodeXGlueCcCodeCompletionLine23config_names:24- go25- java26- javascript27- php28- python29- ruby30dataset_info:31- config_name: java32  features:33  - name: id34    dtype: int3235  - name: input36    dtype: string37  - name: gt38    dtype: string39  splits:40  - name: train41    num_bytes: 545477542    num_examples: 300043  download_size: 169667944  dataset_size: 545477545- config_name: python46  features:47  - name: id48    dtype: int3249  - name: input50    dtype: string51  - name: gt52    dtype: string53  splits:54  - name: train55    num_bytes: 2402155456    num_examples: 1000057  download_size: 814067058  dataset_size: 2402155459configs:60- config_name: java61  data_files:62  - split: train63    path: java/train-*64- config_name: python65  data_files:66  - split: train67    path: python/train-*68---69# Dataset Card for "code_x_glue_cc_code_completion_line"70 71## Table of Contents72- [Dataset Description](#dataset-description)73  - [Dataset Summary](#dataset-summary)74  - [Supported Tasks and Leaderboards](#supported-tasks)75  - [Languages](#languages)76- [Dataset Structure](#dataset-structure)77  - [Data Instances](#data-instances)78  - [Data Fields](#data-fields)79  - [Data Splits](#data-splits-sample-size)80- [Dataset Creation](#dataset-creation)81  - [Curation Rationale](#curation-rationale)82  - [Source Data](#source-data)83  - [Annotations](#annotations)84  - [Personal and Sensitive Information](#personal-and-sensitive-information)85- [Considerations for Using the Data](#considerations-for-using-the-data)86  - [Social Impact of Dataset](#social-impact-of-dataset)87  - [Discussion of Biases](#discussion-of-biases)88  - [Other Known Limitations](#other-known-limitations)89- [Additional Information](#additional-information)90  - [Dataset Curators](#dataset-curators)91  - [Licensing Information](#licensing-information)92  - [Citation Information](#citation-information)93  - [Contributions](#contributions)94 95## Dataset Description96 97- **Homepage:** https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line98 99### Dataset Summary100 101CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line102 103Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity.104We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but fail in completing an unfinished line like a method call with specific parameters, a function signature, a loop condition, a variable definition and so on. When a software develop finish one or more tokens of the current line, the line level completion model is expected to generate the entire line of syntactically correct code.105Line level code completion task shares the train/dev dataset with token level completion. After training a model on CodeCompletion-token, you could directly use it to test on line-level completion.106 107### Supported Tasks and Leaderboards108 109- `slot-filling`: The dataset can be used to train a model for completing entire code lines.110 111### Languages112 113- Java **programming** language114- Python **programming** language115 116## Dataset Structure117 118### Data Instances119 120#### java121 122An example of 'train' looks as follows.123```124{125    "gt": "", 126    "id": 0, 127    "input": "<s> package org . rubypeople . rdt . internal . ui . rubyeditor ; import java . util . Iterator ; import org . eclipse . core . resources . IMarker ; import org . eclipse . ui . texteditor . MarkerAnnotation ; import org . eclipse . ui . texteditor . MarkerUtilities ; import org . rubypeople . rdt . core . IRubyElement ; import org . rubypeople . rdt . core . IRubyModelMarker ; import org . rubypeople . rdt . core . IRubyScript ; import org . rubypeople . rdt . core . RubyCore ; public class RubyMarkerAnnotation extends MarkerAnnotation implements IRubyAnnotation { public static final String RUBY_MARKER_TYPE_PREFIX = \"\" ; public static final String ERROR_ANNOTATION_TYPE = \"\" ; public static final String WARNING_ANNOTATION_TYPE = \"\" ; public static final String INFO_ANNOTATION_TYPE = \"\" ; public static final String TASK_ANNOTATION_TYPE = \"\" ; private IRubyAnnotation fOverlay ; public RubyMarkerAnnotation ( IMarker marker ) { super ( marker ) ; } public String [ ] getArguments ( ) { return null ; } public int getId ( ) { IMarker marker = getMarker ( ) ; if ( marker == null || ! marker . exists ( ) ) return - 1 ; if ( isProblem ( ) ) return marker . getAttribute ( IRubyModelMarker . ID , - 1 ) ; return - 1 ; } public boolean isProblem ( ) { String type = getType ( ) ; return WARNING_ANNOTATION_TYPE . equals ( type ) || ERROR_ANNOTATION_TYPE . equals"128}129```130 131#### python132 133An example of 'train' looks as follows.134```135{136    "gt": "", 137    "id": 0, 138    "input": "<s> from __future__ import absolute_import <EOL> import weakref <EOL> import operator <EOL> from . compat import threading , itertools_filterfalse <EOL> from . import py2k <EOL> import types <EOL> EMPTY_SET = frozenset ( ) <EOL> class KeyedTuple ( tuple ) : <EOL> def __new__ ( cls , vals , labels = None ) : <EOL> t = tuple . __new__ ( cls , vals ) <EOL> t . _labels = [ ] <EOL> if labels : <EOL> t . __dict__ . update ( zip ( labels , vals ) ) <EOL> t . _labels = labels <EOL> return t <EOL> def keys ( self ) : <EOL> return [ l for l in self . _labels if l is not None ] <EOL> @ property <EOL> def _fields ( self ) : <EOL> return tuple ( self . keys ( ) ) <EOL> def _asdict ( self ) : <EOL> return dict ( ( key , self . __dict__ [ key ] ) for key in self . keys ( ) ) <EOL> class ImmutableContainer ( object ) : <EOL> def _immutable ( self , * arg , ** kw ) : <EOL> raise TypeError ( \"\" % self . __class__ . __name__ ) <EOL> __delitem__ = __setitem__ = __setattr__ = _immutable <EOL> class immutabledict ( ImmutableContainer , dict ) : <EOL> clear = pop = popitem = setdefault = update = ImmutableContainer . _immutable <EOL> def __new__ ( cls , * args ) : <EOL> new = dict . __new__ ( cls ) <EOL> dict . __init__ ( new , * args ) <EOL> return new <EOL> def __init__ ( self , * args ) : <EOL> pass <EOL> def __reduce__ ( self ) : <EOL> return immutabledict , ( dict ( self ) , ) <EOL> def union ( self , d ) : <EOL> if not self : <EOL> return immutabledict ( d ) <EOL> else : <EOL> d2 = immutabledict ( self ) <EOL> dict . update ( d2 , d ) <EOL> return d2 <EOL> def __repr__ ( self ) : <EOL> return \"\" % dict . __repr__ ( self ) <EOL> class Properties ( object ) : <EOL> def __init__ ( self , data ) : <EOL> self . __dict__ [ '_data' ] = data <EOL> def __len__ ( self ) : <EOL> return len ( self . _data ) <EOL> def __iter__ ( self ) : <EOL> return iter ( list ( self . _data . values ( ) ) ) <EOL> def __add__ ( self , other ) : <EOL> return list ( self ) + list ( other ) <EOL> def __setitem__ ( self , key , object ) : <EOL> self . _data [ key ] = object <EOL> def __getitem__ ( self , key ) : <EOL> return self . _data [ key ] <EOL> def __delitem__ ( self , key ) : <EOL> del self . _data [ key ] <EOL> def __setattr__ ( self , key , object ) : <EOL> self . _data [ key ] = object <EOL> def __getstate__ ( self ) : <EOL> return { '_data' : self . __dict__ [ '_data' ] } <EOL> def __setstate__ ( self , state ) : <EOL> self . __dict__ [ '_data' ] = state [ '_data' ] <EOL> def __getattr__ ( self , key ) : <EOL> try : <EOL> return self . _data [ key ] <EOL> except KeyError : <EOL> raise AttributeError ( key ) <EOL> def __contains__ ( self , key ) : <EOL> return key in self . _data <EOL> def as_immutable ( self ) : <EOL> return ImmutableProperties ( self . _data ) <EOL> def update ( self , value ) : <EOL> self . _data . update ( value ) <EOL> def get ( self , key , default = None ) : <EOL> if key in self : <EOL> return self [ key ] <EOL> else : <EOL> return default <EOL> def keys ( self ) : <EOL> return list ( self . _data ) <EOL> def values ( self ) : <EOL> return list ( self . _data . values ( ) ) <EOL> def items ( self ) : <EOL> return list ( self . _data . items ( ) ) <EOL> def has_key ( self , key ) : <EOL> return key in self . _data <EOL> def clear ( self ) : <EOL> self . _data . clear ( ) <EOL> class OrderedProperties ( Properties ) : <EOL> def __init__ ( self ) : <EOL> Properties . __init__ ( self , OrderedDict ( ) ) <EOL> class ImmutableProperties ( ImmutableContainer , Properties ) : <EOL> class OrderedDict ( dict ) : <EOL> def __init__ ( self , ____sequence = None , ** kwargs ) : <EOL> self . _list = [ ] <EOL> if ____sequence is None : <EOL> if kwargs : <EOL> self . update ( ** kwargs ) <EOL> else : <EOL> self . update ( ____sequence , ** kwargs ) <EOL> def clear ( self ) : <EOL> self . _list = [ ] <EOL> dict . clear ( self ) <EOL> def copy ( self ) : <EOL> return self . __copy__ ( ) <EOL> def __copy__ ( self ) : <EOL> return OrderedDict ( self ) <EOL> def sort ( self , * arg , ** kw ) : <EOL> self . _list . sort ( * arg , ** kw ) <EOL> def update ( self , ____sequence = None , ** kwargs ) : <EOL> if ____sequence is not None : <EOL> if hasattr ( ____sequence , 'keys' ) : <EOL> for key in ____sequence . keys ( ) : <EOL> self . __setitem__ ( key , ____sequence [ key ] ) <EOL> else : <EOL> for key , value in ____sequence : <EOL> self [ key ] = value <EOL> if kwargs : <EOL> self . update ( kwargs ) <EOL> def setdefault ( self , key , value ) : <EOL> if key not in self : <EOL> self . __setitem__ ( key , value ) <EOL> return value <EOL> else : <EOL> return self . __getitem__ ( key ) <EOL> def __iter__ ( self ) : <EOL> return iter ( self . _list ) <EOL> def keys ( self ) : <EOL> return list ( self ) <EOL> def values ( self ) : <EOL> return [ self [ key ] for key in self . _list ] <EOL> def items ( self ) : <EOL> return [ ( key , self [ key ] ) for key in self . _list ] <EOL> if py2k : <EOL> def itervalues ( self ) : <EOL> return iter ( self . values ( ) ) <EOL> def iterkeys ( self ) : <EOL> return iter ( self ) <EOL> def iteritems ( self ) : <EOL> return iter ( self . items ( ) ) <EOL> def __setitem__ ( self , key , object ) : <EOL> if key not in self : <EOL> try : <EOL> self . _list . append ( key ) <EOL> except AttributeError : <EOL> self . _list = [ key ] <EOL> dict . __setitem__ ( self , key , object ) <EOL> def __delitem__ ( self , key ) : <EOL> dict . __delitem__ ( self , key ) <EOL> self . _list . remove ( key ) <EOL> def pop ( self , key , * default ) : <EOL> present = key in self <EOL> value = dict . pop ( self , key , * default ) <EOL> if present : <EOL> self . _list . remove ( key ) <EOL> return value <EOL> def popitem ( self ) : <EOL> item = dict . popitem ( self ) <EOL> self . _list . remove ( item [ 0 ] ) <EOL> return item <EOL> class OrderedSet ( set ) : <EOL> def __init__ ( self , d = None ) : <EOL> set . __init__ ( self ) <EOL> self . _list = [ ] <EOL> if d is not None : <EOL>"139}140```141 142### Data Fields143 144In the following each data field in go is explained for each config. The data fields are the same among all splits.145 146#### java, python147 148|field name| type |        description         |149|----------|------|----------------------------|150|id        |int32 | Index of the sample        |151|input     |string| Input code string          |152|gt        |string| Code string to be predicted|153 154### Data Splits155 156| name |train|157|------|----:|158|java  | 3000|159|python|10000|160 161## Dataset Creation162 163### Curation Rationale164 165[More Information Needed]166 167### Source Data168 169#### Initial Data Collection and Normalization170 171[More Information Needed]172 173#### Who are the source language producers?174 175[More Information Needed]176 177### Annotations178 179#### Annotation process180 181[More Information Needed]182 183#### Who are the annotators?184 185[More Information Needed]186 187### Personal and Sensitive Information188 189[More Information Needed]190 191## Considerations for Using the Data192 193### Social Impact of Dataset194 195[More Information Needed]196 197### Discussion of Biases198 199[More Information Needed]200 201### Other Known Limitations202 203[More Information Needed]204 205## Additional Information206 207### Dataset Curators208 209https://github.com/microsoft, https://github.com/madlag210 211### Licensing Information212 213Computational Use of Data Agreement (C-UDA) License.214 215### Citation Information216 217```218@article{raychev2016probabilistic,219  title={Probabilistic Model for Code with Decision Trees},220  author={Raychev, Veselin and Bielik, Pavol and Vechev, Martin},221  journal={ACM SIGPLAN Notices},222  pages={731--747},223  year={2016},224  publisher={ACM New York, NY, USA}225}226@inproceedings{allamanis2013mining,227  title={Mining Source Code Repositories at Massive Scale using Language Modeling},228  author={Allamanis, Miltiadis and Sutton, Charles},229  booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)},230  pages={207--216},231  year={2013},232  organization={IEEE}233}234```235 236### Contributions237 238Thanks to @madlag (and partly also @ncoop57) for adding this dataset.