izhx/google-code-jam
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others.
036
This dataset is created by DeepSim: deep learning code functional similarity.
I downloaded googlejam4.tar.gz from parasol-aser/deepsim, fixed encoding of 6/googlejam6.p261.Round1B.java and 1/googlejam1.p815.MushroomMonster.java, and re-compressed.
The all split (all 12 problems) is consistent with their paper (I guess...).
The test split (problem 5, 6, 7, 8, 12) is used in experiments of Language Models are Universal Embedders.
