google/code_x_glue_cc_code_completion_line
Dataset Card for "code_x_glue_cc_code_completion_line" Dataset Summary CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity. We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_line.
10457
1---2annotations_creators:3- found4language_creators:5- found6language:7- code8license:9- c-uda10multilinguality:11- monolingual12size_categories:13- 1K<n<10K14- n<1K15source_datasets:16- original17task_categories:18- text-generation19- fill-mask20task_ids:21- slot-filling22pretty_name: CodeXGlueCcCodeCompletionLine23config_names:24- go25- java26- javascript27- php28- python29- ruby30dataset_info:31- config_name: java32 features:33 - name: id34 dtype: int3235 - name: input36 dtype: string37 - name: gt38 dtype: string39 splits:40 - name: train41 num_bytes: 545477542 num_examples: 300043 download_size: 169667944 dataset_size: 545477545- config_name: python46 features:47 - name: id48 dtype: int3249 - name: input50 dtype: string51 - name: gt52 dtype: string53 splits:54 - name: train55 num_bytes: 2402155456 num_examples: 1000057 download_size: 814067058 dataset_size: 2402155459configs:60- config_name: java61 data_files:62 - split: train63 path: java/train-*64- config_name: python65 data_files:66 - split: train67 path: python/train-*68---69# Dataset Card for "code_x_glue_cc_code_completion_line"70 71## Table of Contents72- [Dataset Description](#dataset-description)73 - [Dataset Summary](#dataset-summary)74 - [Supported Tasks and Leaderboards](#supported-tasks)75 - [Languages](#languages)76- [Dataset Structure](#dataset-structure)77 - [Data Instances](#data-instances)78 - [Data Fields](#data-fields)79 - [Data Splits](#data-splits-sample-size)80- [Dataset Creation](#dataset-creation)81 - [Curation Rationale](#curation-rationale)82 - [Source Data](#source-data)83 - [Annotations](#annotations)84 - [Personal and Sensitive Information](#personal-and-sensitive-information)85- [Considerations for Using the Data](#considerations-for-using-the-data)86 - [Social Impact of Dataset](#social-impact-of-dataset)87 - [Discussion of Biases](#discussion-of-biases)88 - [Other Known Limitations](#other-known-limitations)89- [Additional Information](#additional-information)90 - [Dataset Curators](#dataset-curators)91 - [Licensing Information](#licensing-information)92 - [Citation Information](#citation-information)93 - [Contributions](#contributions)94 95## Dataset Description96 97- **Homepage:** https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line98 99### Dataset Summary100 101CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line102 103Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity.104We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but fail in completing an unfinished line like a method call with specific parameters, a function signature, a loop condition, a variable definition and so on. When a software develop finish one or more tokens of the current line, the line level completion model is expected to generate the entire line of syntactically correct code.105Line level code completion task shares the train/dev dataset with token level completion. After training a model on CodeCompletion-token, you could directly use it to test on line-level completion.106 107### Supported Tasks and Leaderboards108 109- `slot-filling`: The dataset can be used to train a model for completing entire code lines.110 111### Languages112 113- Java **programming** language114- Python **programming** language115 116## Dataset Structure117 118### Data Instances119 120#### java121 122An example of 'train' looks as follows.123```124{125 "gt": "", 126 "id": 0, 127 "input": "<s> package org . rubypeople . rdt . internal . ui . rubyeditor ; import java . util . Iterator ; import org . eclipse . core . resources . IMarker ; import org . eclipse . ui . texteditor . MarkerAnnotation ; import org . eclipse . ui . texteditor . MarkerUtilities ; import org . rubypeople . rdt . core . IRubyElement ; import org . rubypeople . rdt . core . IRubyModelMarker ; import org . rubypeople . rdt . core . IRubyScript ; import org . rubypeople . rdt . core . RubyCore ; public class RubyMarkerAnnotation extends MarkerAnnotation implements IRubyAnnotation { public static final String RUBY_MARKER_TYPE_PREFIX = \"\" ; public static final String ERROR_ANNOTATION_TYPE = \"\" ; public static final String WARNING_ANNOTATION_TYPE = \"\" ; public static final String INFO_ANNOTATION_TYPE = \"\" ; public static final String TASK_ANNOTATION_TYPE = \"\" ; private IRubyAnnotation fOverlay ; public RubyMarkerAnnotation ( IMarker marker ) { super ( marker ) ; } public String [ ] getArguments ( ) { return null ; } public int getId ( ) { IMarker marker = getMarker ( ) ; if ( marker == null || ! marker . exists ( ) ) return - 1 ; if ( isProblem ( ) ) return marker . getAttribute ( IRubyModelMarker . ID , - 1 ) ; return - 1 ; } public boolean isProblem ( ) { String type = getType ( ) ; return WARNING_ANNOTATION_TYPE . equals ( type ) || ERROR_ANNOTATION_TYPE . equals"128}129```130 131#### python132 133An example of 'train' looks as follows.134```135{136 "gt": "", 137 "id": 0, 138 "input": "<s> from __future__ import absolute_import <EOL> import weakref <EOL> import operator <EOL> from . compat import threading , itertools_filterfalse <EOL> from . import py2k <EOL> import types <EOL> EMPTY_SET = frozenset ( ) <EOL> class KeyedTuple ( tuple ) : <EOL> def __new__ ( cls , vals , labels = None ) : <EOL> t = tuple . __new__ ( cls , vals ) <EOL> t . _labels = [ ] <EOL> if labels : <EOL> t . __dict__ . update ( zip ( labels , vals ) ) <EOL> t . _labels = labels <EOL> return t <EOL> def keys ( self ) : <EOL> return [ l for l in self . _labels if l is not None ] <EOL> @ property <EOL> def _fields ( self ) : <EOL> return tuple ( self . keys ( ) ) <EOL> def _asdict ( self ) : <EOL> return dict ( ( key , self . __dict__ [ key ] ) for key in self . keys ( ) ) <EOL> class ImmutableContainer ( object ) : <EOL> def _immutable ( self , * arg , ** kw ) : <EOL> raise TypeError ( \"\" % self . __class__ . __name__ ) <EOL> __delitem__ = __setitem__ = __setattr__ = _immutable <EOL> class immutabledict ( ImmutableContainer , dict ) : <EOL> clear = pop = popitem = setdefault = update = ImmutableContainer . _immutable <EOL> def __new__ ( cls , * args ) : <EOL> new = dict . __new__ ( cls ) <EOL> dict . __init__ ( new , * args ) <EOL> return new <EOL> def __init__ ( self , * args ) : <EOL> pass <EOL> def __reduce__ ( self ) : <EOL> return immutabledict , ( dict ( self ) , ) <EOL> def union ( self , d ) : <EOL> if not self : <EOL> return immutabledict ( d ) <EOL> else : <EOL> d2 = immutabledict ( self ) <EOL> dict . update ( d2 , d ) <EOL> return d2 <EOL> def __repr__ ( self ) : <EOL> return \"\" % dict . __repr__ ( self ) <EOL> class Properties ( object ) : <EOL> def __init__ ( self , data ) : <EOL> self . __dict__ [ '_data' ] = data <EOL> def __len__ ( self ) : <EOL> return len ( self . _data ) <EOL> def __iter__ ( self ) : <EOL> return iter ( list ( self . _data . values ( ) ) ) <EOL> def __add__ ( self , other ) : <EOL> return list ( self ) + list ( other ) <EOL> def __setitem__ ( self , key , object ) : <EOL> self . _data [ key ] = object <EOL> def __getitem__ ( self , key ) : <EOL> return self . _data [ key ] <EOL> def __delitem__ ( self , key ) : <EOL> del self . _data [ key ] <EOL> def __setattr__ ( self , key , object ) : <EOL> self . _data [ key ] = object <EOL> def __getstate__ ( self ) : <EOL> return { '_data' : self . __dict__ [ '_data' ] } <EOL> def __setstate__ ( self , state ) : <EOL> self . __dict__ [ '_data' ] = state [ '_data' ] <EOL> def __getattr__ ( self , key ) : <EOL> try : <EOL> return self . _data [ key ] <EOL> except KeyError : <EOL> raise AttributeError ( key ) <EOL> def __contains__ ( self , key ) : <EOL> return key in self . _data <EOL> def as_immutable ( self ) : <EOL> return ImmutableProperties ( self . _data ) <EOL> def update ( self , value ) : <EOL> self . _data . update ( value ) <EOL> def get ( self , key , default = None ) : <EOL> if key in self : <EOL> return self [ key ] <EOL> else : <EOL> return default <EOL> def keys ( self ) : <EOL> return list ( self . _data ) <EOL> def values ( self ) : <EOL> return list ( self . _data . values ( ) ) <EOL> def items ( self ) : <EOL> return list ( self . _data . items ( ) ) <EOL> def has_key ( self , key ) : <EOL> return key in self . _data <EOL> def clear ( self ) : <EOL> self . _data . clear ( ) <EOL> class OrderedProperties ( Properties ) : <EOL> def __init__ ( self ) : <EOL> Properties . __init__ ( self , OrderedDict ( ) ) <EOL> class ImmutableProperties ( ImmutableContainer , Properties ) : <EOL> class OrderedDict ( dict ) : <EOL> def __init__ ( self , ____sequence = None , ** kwargs ) : <EOL> self . _list = [ ] <EOL> if ____sequence is None : <EOL> if kwargs : <EOL> self . update ( ** kwargs ) <EOL> else : <EOL> self . update ( ____sequence , ** kwargs ) <EOL> def clear ( self ) : <EOL> self . _list = [ ] <EOL> dict . clear ( self ) <EOL> def copy ( self ) : <EOL> return self . __copy__ ( ) <EOL> def __copy__ ( self ) : <EOL> return OrderedDict ( self ) <EOL> def sort ( self , * arg , ** kw ) : <EOL> self . _list . sort ( * arg , ** kw ) <EOL> def update ( self , ____sequence = None , ** kwargs ) : <EOL> if ____sequence is not None : <EOL> if hasattr ( ____sequence , 'keys' ) : <EOL> for key in ____sequence . keys ( ) : <EOL> self . __setitem__ ( key , ____sequence [ key ] ) <EOL> else : <EOL> for key , value in ____sequence : <EOL> self [ key ] = value <EOL> if kwargs : <EOL> self . update ( kwargs ) <EOL> def setdefault ( self , key , value ) : <EOL> if key not in self : <EOL> self . __setitem__ ( key , value ) <EOL> return value <EOL> else : <EOL> return self . __getitem__ ( key ) <EOL> def __iter__ ( self ) : <EOL> return iter ( self . _list ) <EOL> def keys ( self ) : <EOL> return list ( self ) <EOL> def values ( self ) : <EOL> return [ self [ key ] for key in self . _list ] <EOL> def items ( self ) : <EOL> return [ ( key , self [ key ] ) for key in self . _list ] <EOL> if py2k : <EOL> def itervalues ( self ) : <EOL> return iter ( self . values ( ) ) <EOL> def iterkeys ( self ) : <EOL> return iter ( self ) <EOL> def iteritems ( self ) : <EOL> return iter ( self . items ( ) ) <EOL> def __setitem__ ( self , key , object ) : <EOL> if key not in self : <EOL> try : <EOL> self . _list . append ( key ) <EOL> except AttributeError : <EOL> self . _list = [ key ] <EOL> dict . __setitem__ ( self , key , object ) <EOL> def __delitem__ ( self , key ) : <EOL> dict . __delitem__ ( self , key ) <EOL> self . _list . remove ( key ) <EOL> def pop ( self , key , * default ) : <EOL> present = key in self <EOL> value = dict . pop ( self , key , * default ) <EOL> if present : <EOL> self . _list . remove ( key ) <EOL> return value <EOL> def popitem ( self ) : <EOL> item = dict . popitem ( self ) <EOL> self . _list . remove ( item [ 0 ] ) <EOL> return item <EOL> class OrderedSet ( set ) : <EOL> def __init__ ( self , d = None ) : <EOL> set . __init__ ( self ) <EOL> self . _list = [ ] <EOL> if d is not None : <EOL>"139}140```141 142### Data Fields143 144In the following each data field in go is explained for each config. The data fields are the same among all splits.145 146#### java, python147 148|field name| type | description |149|----------|------|----------------------------|150|id |int32 | Index of the sample |151|input |string| Input code string |152|gt |string| Code string to be predicted|153 154### Data Splits155 156| name |train|157|------|----:|158|java | 3000|159|python|10000|160 161## Dataset Creation162 163### Curation Rationale164 165[More Information Needed]166 167### Source Data168 169#### Initial Data Collection and Normalization170 171[More Information Needed]172 173#### Who are the source language producers?174 175[More Information Needed]176 177### Annotations178 179#### Annotation process180 181[More Information Needed]182 183#### Who are the annotators?184 185[More Information Needed]186 187### Personal and Sensitive Information188 189[More Information Needed]190 191## Considerations for Using the Data192 193### Social Impact of Dataset194 195[More Information Needed]196 197### Discussion of Biases198 199[More Information Needed]200 201### Other Known Limitations202 203[More Information Needed]204 205## Additional Information206 207### Dataset Curators208 209https://github.com/microsoft, https://github.com/madlag210 211### Licensing Information212 213Computational Use of Data Agreement (C-UDA) License.214 215### Citation Information216 217```218@article{raychev2016probabilistic,219 title={Probabilistic Model for Code with Decision Trees},220 author={Raychev, Veselin and Bielik, Pavol and Vechev, Martin},221 journal={ACM SIGPLAN Notices},222 pages={731--747},223 year={2016},224 publisher={ACM New York, NY, USA}225}226@inproceedings{allamanis2013mining,227 title={Mining Source Code Repositories at Massive Scale using Language Modeling},228 author={Allamanis, Miltiadis and Sutton, Charles},229 booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)},230 pages={207--216},231 year={2013},232 organization={IEEE}233}234```235 236### Contributions237 238Thanks to @madlag (and partly also @ncoop57) for adding this dataset.