Team Ai
Modelpublic

codesage/codesage-base

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
4likes88downloads
README.md66 linesDownload Raw Back to root
1---2license: apache-2.03datasets:4- bigcode/the-stack-dedup5library_name: transformers6language:7- code8---9 10## CodeSage-Base11 12### Updates13* [12/2024] <span style="color:blue">We are excited to announce the release of the CodeSage V2 model family with largely improved performance and flexible embedding dimensions!</span> Please check out our [models](https://huggingface.co/codesage) and [blogpost](https://code-representation-learning.github.io/codesage-v2.html) for more details.14* [11/2024] You can now access CodeSage models through SentenceTransformer.15 16 17### Model description18CodeSage is a new family of open code embedding models with an encoder architecture that support a wide range of source code understanding tasks. It is introduced in the paper:19 20[Code Representation Learning At Scale by 21Dejiao Zhang*, Wasi Uddin Ahmad*, Ming Tan, Hantian Ding, Ramesh Nallapati, Dan Roth, Xiaofei Ma, Bing Xiang](https://arxiv.org/abs/2402.01935) (* indicates equal contribution).22 23### Pretraining data24This checkpoint is trained on the Stack data (https://huggingface.co/datasets/bigcode/the-stack-dedup). Supported languages (9 in total) are as follows: c, c-sharp, go, java, javascript, typescript, php, python, ruby.25 26### Training procedure27This checkpoint is first trained on code data via masked language modeling (MLM) and then on bimodal text-code pair data. Please refer to the paper for more details.28 29### How to Use30This checkpoint consists of an encoder (356M model), which can be used to extract code embeddings of 1024 dimension. 31 321. Accessing CodeSage via HuggingFace: it can be easily loaded using the AutoModel functionality and employs the [Starcoder Tokenizer](https://arxiv.org/pdf/2305.06161.pdf).33   34```35from transformers import AutoModel, AutoTokenizer36 37checkpoint = "codesage/codesage-base"38device = "cuda"  # "cpu" for CPU usage39 40# Note: CodeSage requires adding eos token at the end of each tokenized sequence 41 42tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True, add_eos_token=True)43 44model = AutoModel.from_pretrained(checkpoint, trust_remote_code=True).to(device)45 46inputs = tokenizer.encode("def print_hello_world():\tprint('Hello World!')", return_tensors="pt").to(device)47embedding = model(inputs)[0]48```49 502. Accessing CodeSage via SentenceTransformer51```52from sentence_transformers import SentenceTransformer53model = SentenceTransformer("codesage/codesage-base", trust_remote_code=True)54```55 56### BibTeX entry and citation info57```58@inproceedings{59    zhang2024codesage,60    title={CodeSage: Code Representation Learning At Scale},61    author={Dejiao Zhang* and Wasi Ahmad* and Ming Tan and Hantian Ding and Ramesh Nallapati and Dan Roth and Xiaofei Ma and Bing Xiang},62    booktitle={The Twelfth International Conference on Learning Representations},63    year={2024},64    url={https://openreview.net/forum?id=vfzRRjumpX}65}66```