JetBrains-Research/lca-codegen-small
LCA Project Level Code Completion How to load the dataset from datasets import load_dataset ds = load_dataset('JetBrains-Research/lca-codegen-small', split='test') Data Point Structure repo – repository name in format {GitHub_user_name}__{repository_name} commit_hash – commit hash completion_file – dictionary with the completion file content in the following format: filename – filepath to the completion file content – content of the completion… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-codegen-small.
LCA Project Level Code Completion
How to load the dataset
from datasets import load_dataset
ds = load_dataset('JetBrains-Research/lca-codegen-small', split='test')Data Point Structure
repo– repository name in format{GitHub_user_name}__{repository_name}commit_hash– commit hashcompletion_file– dictionary with the completion file content in the following format:filename– filepath to the completion filecontent– content of the completion filecompletion_lines– dictionary where keys are classes of lines and values are a list of integers (numbers of lines to complete). The classes are:committed– line contains at least one function or class that was declared in the committed files fromcommit_hashinproject– line contains at least one function or class that was declared in the project (excluding previous)infile– line contains at least one function or class that was declared in the completion file (excluding previous)common– line contains at least one function or class that was classified to be common, e.g.,main,get, etc (excluding previous)non_informative– line that was classified to be non-informative, e.g. too short, contains comments, etcrandom– randomly sampled from the rest of the linesrepo_snapshot– dictionary with a snapshot of the repository before the commit. Has the same structure ascompletion_file, but filenames and contents are organized as lists.completion_lines_raw– the same ascompletion_lines, but before sampling.
How we collected the data
To collect the data, we cloned repositories from GitHub where the main language is Python. The completion file for each data point is a .py file that was added to the repository in a commit. The state of the repository before this commit is the repo snapshot.
Small dataset is defined by number of characters in .py files from the repository snapshot. This number is less than 48K.
Dataset Stats
- Number of datapoints: 144
- Number of repositories: 46
- Number of commits: 63
Completion File
- Number of lines, median: 310.5
- Number of lines, min: 201
- Number of lines, max: 1916
Repository Snapshot
.pyfiles: <u>median 4</u>, from 0 to 52- non
.pyfiles: <u>median 19.5</u>, from 1 to 1044 .pylines: <u>median 128</u>- non
.pylines: <u>median 1227</u>
Line Counts:
- infile: 1430
- inproject: 95
- common: 500
- committed: 1426
- non-informative: 532
- random: 703
- total: 4686
