Team Ai
Modelpublic

nomic-ai/CodeRankEmbed

sourceHugging Facemitupdated 1y agoView on Hugging Face
79likes162kdownloads
README.md69 linesDownload Raw Back to root
1---2base_model:3- Snowflake/snowflake-arctic-embed-m-long4library_name: sentence-transformers5license: mit6---7 8 9# CodeRankEmbed10 11`CodeRankEmbed` is a 137M bi-encoder supporting 8192 context length for code retrieval. It significantly outperforms various open-source and proprietary code embedding models on various code retrieval tasks.  12 13Check out our [blog post](https://gangiswag.github.io/cornstack/) and [paper](https://arxiv.org/pdf/2412.01007) for more details!14 15Combine `CodeRankEmbed` with our re-ranker [`CodeRankLLM`](https://huggingface.co/cornstack/CodeRankLLM) for even higher quality code retrieval.16 17# Performance Benchmarks18 19| Name                             | Parameters | CSN (MRR)      | CoIR (NDCG@10)     |20| :-------------------------------:| :----- | :-------- | :------: | 21| **CodeRankEmbed**              | 137M   | **77.9** |**60.1** | 22| Arctic-Embed-M-Long       | 137M   | 53.4    | 43.0    | 23| CodeSage-Small       | 130M   | 64.9    | 54.4    | 24| CodeSage-Base       | 356M   | 68.7    | 57.5    | 25| CodeSage-Large       | 1.3B   | 71.2    | 59.4    | 26| Jina-Code-v2           | 161M   | 67.2     | 58.4  |27| CodeT5+          | 110M   | 74.2     | 45.9     | 28| OpenAI-Ada-002          | 110M   | 71.3     | 45.6     | 29| Voyage-Code-002        | Unknown   | 68.5     | 56.3     | 30 31 32We release the scripts to evaluate our model's performance [here](https://github.com/gangiswag/cornstack).33 34# Usage35 36**Important**: the query prompt *must* include the following *task instruction prefix*: "Represent this query for searching relevant code" 37 38```python39from sentence_transformers import SentenceTransformer40 41model = SentenceTransformer("nomic-ai/CodeRankEmbed", trust_remote_code=True)42queries = ['Represent this query for searching relevant code: Calculate the n-th factorial']43codes = ['def fact(n):\n if n < 0:\n  raise ValueError\n return 1 if n == 0 else n * fact(n - 1)']44query_embeddings = model.encode(queries)45print(query_embeddings)46code_embeddings = model.encode(codes)47print(code_embeddings)48```49 50 51 52## Training53We use a bi-encoder architecture for `CodeRankEmbed`, with weights shared between the text and code encoder. The retriever is contrastively fine-tuned with InfoNCE loss on a 21 million example high-quality dataset we curated called [CoRNStack](https://gangiswag.github.io/cornstack/). Our encoder is initialized with [Arctic-Embed-M-Long](https://huggingface.co/Snowflake/snowflake-arctic-embed-m-long), a 137M parameter text encoder supporting an extended context length of 8,192 tokens.54 55# Citation56 57If you find the model, dataset, or training code useful, please cite our work:58 59```bibtex60@misc{suresh2025cornstackhighqualitycontrastivedata,61      title={CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking}, 62      author={Tarun Suresh and Revanth Gangi Reddy and Yifei Xu and Zach Nussbaum and Andriy Mulyar and Brandon Duderstadt and Heng Ji},63      year={2025},64      eprint={2412.01007},65      archivePrefix={arXiv},66      primaryClass={cs.CL},67      url={https://arxiv.org/abs/2412.01007}, 68}69```