Team Ai
Datasetpublic

nomic-ai/cornstack-ruby-v1

CoRNStack Ruby Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-ruby-v1.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes270downloads
4 commits on main
29222462y ago

Update README.md

zpn
ab76f622y ago

Create README.md

zpn
ef968382y ago

Upload folder using huggingface_hub

gangiswag
99447432y ago

initial commit

gangiswag