ChrisGVE/codebert-base-Q8_0-GGUF
077
CodeBERT base — GGUF Q8_0
`microsoft/codebert-base` converted to GGUF at Q8_0 for llama.cpp's llama-server --embeddings.
llama-server -m codebert-base-Q8_0.gguf --embeddings --port 8080
curl -s localhost:8080/v1/embeddings -H 'Content-Type: application/json' -d '{"input":["def f(): pass"]}'How it was converted
convert_hf_to_gguf.py (llama.cpp, 2026-10-01) with --outtype q8_0, after two additions to the source directory:
- a
tokenizer.json, written byRobertaTokenizerFast.save_pretrained. The upstream repo ships onlyvocab.json+merges.txt, and withouttokenizer.jsonthe converter falls back to a WordPiece vocabulary and tokenizes code wrongly; - a sentence-transformers
modules.json+1_Pooling/config.jsondeclaring mean pooling, so the pooling type is written into the GGUF.
Verification
40 random code/console/output blocks (first 600 characters), embedded by this file on llama-server (CPU) and by the original model in PyTorch fp32 with masked mean pooling over the last hidden state. Token ids are identical; cosine similarity min 0.9998, median 0.99994.
The weights and their licence (MIT) are Microsoft's; see the CodeBERT paper.
