Team Ai
Datasetpublic

google-research-datasets/great_code

The dataset for the variable-misuse task, described in the ICLR 2020 paper 'Global Relational Models of Source Code' [https://openreview.net/forum?id=B1lnbRNtwr] This is the public version of the dataset used in that paper. The original, used to produce the graphs in the paper, could not be open-sourced due to licensing issues. See the public associated code repository [https://github.com/VHellendoorn/ICLR20-Great] for results produced from this dataset. This dataset was generated synthetically from the corpus of Python code in the ETH Py150 Open dataset [https://github.com/google-research-datasets/eth_py150_open].

sourceHugging Facecc-by-sa-3.0updated 3y agoView on Hugging Face
5likes413downloads
12 commits on main
8e2d62a3y ago

Delete legacy JSON metadata (#2)

albertvillanova
67caa224y ago

Reorder split names (#1)

albertvillanova
f89204f4y ago

add dataset_info in dataset metadata

lhoestq
fccd3da4y ago

remove dummmy data

mariosasko
c750f3d4y ago

Align more metadata with other repo types (models,spaces) (#4607)

julien-c
c556d7a4y ago

Remove config names as yaml keys (#4367)

lhoestq
81a84a34y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
79693645y ago

Update files from the datasets library (from 1.16.0)

system
8740f455y ago

Update files from the datasets library (from 1.7.0)

system
d3e78e35y ago

Update files from the datasets library (from 1.6.0)

system
ddc1e145y ago

Update files from the datasets library (from 1.3.0)

system
73a07275y ago

Update files from the datasets library (from 1.2.0)

system