Team Ai
Modelpublic

nethunter2023/kernel-code-embed

sourceHugging Facegpl-2.0updated 1mo agoView on Hugging Face
0likes38downloads
README.md121 linesDownload Raw Back to root
1---2license: gpl-2.03library_name: sentence-transformers4pipeline_tag: sentence-similarity5tags:6  - sentence-transformers7  - feature-extraction8  - sentence-similarity9  - code-retrieval10  - code-search11  - linux-kernel12  - c13language:14  - code15---16 17# kernel-code-embed18 19A 42.6M-parameter bi-encoder for retrieving **Linux kernel C** from20natural-language queries. Queries and code go through the same encoder, with no21prefix or instruction prompt.22 23512-dim output, mean pooling, L2-normalised — score with cosine similarity.24 25## Usage26 27```python28from sentence_transformers import SentenceTransformer29 30model = SentenceTransformer("nethunter2023/kernel-code-embed")31 32query = "how are free pages coalesced into larger blocks"33code = "static inline void __free_one_page(struct page *page, ...) { ... }"34 35emb = model.encode([query, code], normalize_embeddings=True)36print(emb @ emb.T)          # cosine similarity37```38 39### Sequence length40 41`max_seq_length` is **320**, and should not be raised even though42`max_position_embeddings` is 512. Pretraining saw 512 tokens but the retrieval43stage trained at 320, so positions 320–511 are undertrained; feeding longer44inputs measurably degrades ranking. Chunk longer functions instead.45 46## Results47 48Two different protocols. Read them separately — the candidate pools differ by49two orders of magnitude, so the numbers are **not** comparable across tables.50 51### 1. Open retrieval — the whole kernel52 53Every `.c`/`.h` chunk of Linux v7.1-rc5 as the index (**914,554 candidates**),54**N = 400** held-out queries, one correct answer each.55 56| | this model | + BM25 fusion |57|---|---|---|58| recall@1 | **0.8125** | 0.9050 |59| recall@5 | 0.9375 | 0.9725 |60| recall@10 | 0.9575 | 0.9775 |61| recall@50 | 0.9850 | 0.9950 |62| MRR | 0.8682 | 0.9374 |63 64**The model alone ranks the correct function first 81% of the time out of65914,554 candidates** (95% CI ±3.8pp at N=400).66 67The second column is reciprocal-rank fusion of this model with BM25. It is68higher, but part of that lift is BM25's, so the 0.8125 figure is the one that69belongs to this model.70 71### 2. Closed set — against other encoders72 73**N = 2,000** held-out queries, 4,000 candidates, where each distractor is a74sibling function from the same source file as the answer.75 76| | params | accuracy@1 | NDCG@10 |77|---|---|---|---|78| BM25 (lexical) | — | 0.7115 | 0.8297 |79| `jina-embeddings-v2-base-code` | 161M | 0.9035 | 0.9541 |80| **this model** | **42.6M** | **0.9230** | **0.9661** |81 82Scored with the identical query set, candidate pool and metric code. The83general-purpose code embedder was given a 512-token budget — more than this84model's 320.85 86The +2.0pp accuracy@1 margin over `jina-embeddings-v2-base-code` is significant87under a paired McNemar test: p = 0.003, 95% CI on the paired difference88+0.7pp to +3.4pp (N = 2,000). A 42.6M domain model edging out a 161M89general-purpose one — at 3.8× fewer parameters — is the result worth having.90 91## How the evaluation set was built92 93This matters for reading the numbers above.94 95Queries are **kernel-doc comments written by kernel developers**, not generated96by a language model — so they are not distribution-matched to the training data97by construction. They are held out from training. The **symbol name is stripped98from the query**: left in, `kmalloc_node` would appear in both the query and the99answer's signature and the task would collapse to identifier matching.100 101Training used InfoNCE over in-batch negatives (batch size 24, softmax scale10220.0), with hard negatives drawn as sibling functions from the same file.103 104## Limitations105 106- **Domain-specific.** Linux kernel C only. Not a general-purpose code or text107  embedding model; do not expect transfer to other languages or codebases.108- **Short context** — 320 tokens. Long functions must be chunked.109- **kernel-doc phrasing.** Queries are developer-written documentation, which is110  more precise than typical end-user questions. Expect lower accuracy on casual111  or ambiguous phrasing.112- **Single held-out split**, no seed variance reported.113 114## License115 116GPL-2.0, matching the Linux kernel corpus it was trained on. Whether a GPL117training corpus propagates to model weights is legally unsettled; this is the118conservative reading, chosen deliberately rather than by default. If you need119different terms for commercial use, treat that as an open question to resolve120with counsel rather than an answered one.121