Team Ai
Datasetpublic

izhx/google-code-jam

Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others.

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes36downloads
Dataset Card

This dataset is created by DeepSim: deep learning code functional similarity.

I downloaded googlejam4.tar.gz from parasol-aser/deepsim, fixed encoding of 6/googlejam6.p261.Round1B.java and 1/googlejam1.p815.MushroomMonster.java, and re-compressed.

The all split (all 12 problems) is consistent with their paper (I guess...).

The test split (problem 5, 6, 7, 8, 12) is used in experiments of Language Models are Universal Embedders.

problemcode num
1478
288
3242
438
52
6435
727
8245
968
1018
1120
124