Team Ai
Modelpublic

ChrisGVE/codebert-base-Q8_0-GGUF

sourceHugging Facemitupdated 7d agoView on Hugging Face
0likes77downloads
Model Card

CodeBERT base — GGUF Q8_0

`microsoft/codebert-base` converted to GGUF at Q8_0 for llama.cpp's llama-server --embeddings.

Filecodebert-base-Q8_0.gguf (135 MB)
SHA-25601a85fb726bf4dc063aaac34b33fde08f991b96b3f9f76ff10e28fa2269df5ea
Source revision99d7ef814601faaf7bdc2f774ffa7dade4f4d828 (safetensors)
Context512 tokens
Poolingmean (stored in the file, no --pooling flag needed)
Dimensions768
sh
llama-server -m codebert-base-Q8_0.gguf --embeddings --port 8080
curl -s localhost:8080/v1/embeddings -H 'Content-Type: application/json' -d '{"input":["def f(): pass"]}'

How it was converted

convert_hf_to_gguf.py (llama.cpp, 2026-10-01) with --outtype q8_0, after two additions to the source directory:

  • —a tokenizer.json, written by RobertaTokenizerFast.save_pretrained. The upstream repo ships only vocab.json + merges.txt, and without tokenizer.json the converter falls back to a WordPiece vocabulary and tokenizes code wrongly;
  • —a sentence-transformers modules.json + 1_Pooling/config.json declaring mean pooling, so the pooling type is written into the GGUF.

Verification

40 random code/console/output blocks (first 600 characters), embedded by this file on llama-server (CPU) and by the original model in PyTorch fp32 with masked mean pooling over the last hidden state. Token ids are identical; cosine similarity min 0.9998, median 0.99994.

The weights and their licence (MIT) are Microsoft's; see the CodeBERT paper.