Team Ai
Modelpublic

Shuu12121/NightJar-large-CodeSearch-Embedding

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
0likes17downloads
Model Card

NightJar-large-CodeSearch-Embedding

NightJar-large-CodeSearch-Embedding is a code-retrieval embedding model fine-tuned from `Shuu12121/NightJar-large` (a 346M-parameter ModernBERT encoder pre-trained from scratch on code) with Sentence Transformers. It maps text and code to a 1024-dimensional vector space, so natural-language queries, code snippets and code edits can be compared with cosine similarity.

Use it for natural-language-to-code search, code-to-code retrieval and code-edit search.

Base model`Shuu12121/NightJar-large`
Model typeSentence Transformer (Transformer + [CLS] pooling)
Parameters346M
Output dimension1024
Max sequence length1,024 tokens
Similarity functionCosine similarity
Task prefixnone

Results (MTEB, nDCG@10)

Scores were computed with MTEB 2.5.1. All models below were fine-tuned with the same recipe (see Training).

CodeSearchNetRetrieval

ModelGoJavaJavaScriptPHPPythonRubyAvg
NightJar-large-CodeSearch-Embedding0.96820.93240.83170.89890.95320.88500.9116
NightJar-CodeSearch-Embedding (smaller NightJar)0.96560.92780.82360.89150.94100.86630.9026
NightOwl (same recipe)0.96740.92320.81570.89070.93930.86180.8997

CodeEditSearchRetrieval

ModelCC++GoJavaJSPHPPythonRubyRustScalaShellSwiftTSAvg
NightJar-large-CodeSearch-Embedding0.72200.75370.80620.78310.78460.77430.80480.81600.73960.82480.75710.78820.81450.7822
NightJar-CodeSearch-Embedding (smaller NightJar)0.66870.71050.77160.73650.73860.71620.77110.77230.69000.78590.71190.74700.77590.7382
NightOwl (same recipe)0.66240.71700.78000.75090.75390.73390.77160.78140.71000.77950.72200.74530.78900.7459

NightJar-large-CodeSearch-Embedding scores highest on every language of both benchmarks.

"NightJar-CodeSearch-Embedding" is the smaller `Shuu12121/NightJar` fine-tuned with exactly this setup for the full epoch. "NightOwl (same recipe)" is `Shuu12121/NightOwl` fine-tuned with the same losses and hyperparameters (without the additional-languages dataset).

The fine-tuning data is the same decontaminated set used for NightOwl-CodeEmbedding: overlaps with the CodeSearchNet test splits, and between the commitpackft-derived code-edit data and the CodeEditSearchRetrieval evaluation examples, were removed before training. MTEB names the CodeEditSearchRetrieval evaluation split train, but those examples were not used for fine-tuning.

Usage

bash
pip install -U sentence-transformers
python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding")

sentences = [
    # query
    "Optional. The amount of time that a device will be initially allocated\n"
    "for. This can eventually be extended with the UpdateDeviceSession RPC.\n"
    "Default: 15 minutes.\n\n"
    "Generated from protobuf field <code>.google.protobuf.Duration ttl = 13 "
    "[(.google.api.field_behavior) = OPTIONAL];</code>\n"
    "@return \\Google\\Protobuf\\Duration|null",
    # matching code
    "public function getTtl()\n    {\n        return $this->readOneof(13);\n    }",
    # similar code, wrong field
    "public function getTtl()\n    {\n        return $this->readOneof(7);\n    }",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# (3, 1024)

similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8610, 0.8172],
#         [0.8610, 1.0000, 0.9410],
#         [0.8172, 0.9410, 1.0000]])

The query scores higher against the matching getter (0.861) than against the near-identical one that reads a different field (0.817), even though the two code snippets are very close to each other (0.941).

Input format tips

  • —Pass queries and code as plain strings, with no task prefix. The fine-tuning data used plain text for queries and plain code for documents, without the [NL] / [Code1] markers from pre-training.
  • —The model was trained with sequences up to 1,024 tokens. Longer inputs are truncated.
  • —The model was trained in both directions (query → code and code → query), so it can also be used to find descriptions for code.

Training

The model was fine-tuned for one epoch (3,978 steps, 4,073,472 examples) from Shuu12121/NightJar-large.

Data

Each example is an anchor, a positive, and 15 hard negatives (negative_1 … negative_15). The three datasets were sampled in an 8 : 8 : 1 ratio:

DatasetContent
`Shuu12121/owl_code_search_hard_negative_datasets_V2_kd`docstring → code search, 8 languages
`Shuu12121/owl_code_search_hard_negative_datasets_additional_languages_V2_kd`docstring → code search, 8 more languages
`Shuu12121/codeedit_hard_negative_datasets_kd`code-edit retrieval, 11 languages

Approximate token lengths, from the first 1,000 samples: anchors average 75 tokens, positives 197 tokens, and negatives 188–215 tokens. The maximum is 1,024 tokens for every column.

Loss

`CachedMultipleNegativesRankingLoss` with:

json
{
    "scale": 100.0,
    "similarity_fct": "cos_sim",
    "mini_batch_size": 64,
    "gather_across_devices": false,
    "directions": ["query_to_doc", "doc_to_query"],
    "partition_mode": "joint",
    "hardness_mode": null,
    "hardness_strength": 0.0
}

Hyperparameters

Batch size1,024 (no gradient accumulation)
Learning rate6e-5, cosine schedule, no warmup
Weight decay0.01
Optimizeradamw_torch_fused (β = 0.9 / 0.999, ε = 1e-8)
Epochs1
Precisionbf16
Gradient checkpointingon
Seed42

Training-time evaluation

During training, the model was monitored with Sentence Transformers' InformationRetrievalEvaluator on CodeSearchNet (six languages, validation split). The values below are from the end of training.

MetricGoJavaJavaScriptPHPPythonRubyAvg
Accuracy@10.87400.76300.77300.76600.89000.85100.8195
Recall@100.98100.93500.90300.82200.98900.95700.9312
MRR@100.91840.83090.82190.78980.93020.89450.8643
nDCG@100.93410.85690.84200.79790.94490.91010.8810

These numbers come from a different evaluation set than MTEB, so they are not comparable with the MTEB scores above.

The monitored nDCG@10 rose quickly in the first few hundred steps. After about step 2,000 (roughly half an epoch) it stayed within ±0.01 of its final value for every language, so most of the gain was already reached in the first half of the epoch.

Framework versions

  • —Python 3.10.12
  • —Sentence Transformers 5.3.0
  • —Transformers 4.56.2
  • —PyTorch 2.8.0+cu128
  • —Accelerate 1.12.0
  • —Datasets 3.6.0
  • —Tokenizers 0.22.1

Architecture

text
SentenceTransformer(
  (0): Transformer({'max_seq_length': 1024, 'do_lower_case': False, 'architecture': 'ModernBertModel'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

The encoder is the ModernBERT-based NightJar-large: 28 layers, hidden size 1024, 16 heads, global attention in every 3rd layer and a 128-token sliding window otherwise. The tokenizer keeps indentation and newlines as single tokens. See the NightJar-large model card for details.

Limitations

  • —Retrieval quality was evaluated on CodeSearchNetRetrieval and CodeEditSearchRetrieval only.
  • —Performance varies by language with the amount and quality of the fine-tuning data. Languages not covered by the fine-tuning data are not well supported.
  • —Inputs longer than 1,024 tokens are truncated.
  • —Retrieved code may contain bugs, outdated APIs or insecure patterns. Review it before use.

License

The model weights are released under Apache 2.0. The training datasets carry their own licenses and terms of use. Check them if your use case depends on the provenance of the training data.

Citation

bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
bibtex
@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

日本語版

概要

NightJar-large-CodeSearch-Embeddingは、`Shuu12121/NightJar-large`(コードでゼロから事前学習した、約3.46億パラメータのModernBERTエンコーダー)を、Sentence Transformersでコード検索向けに追加学習した埋め込みモデルです。文章とコードを1024次元のベクトルに変換するので、自然言語の質問、コード断片、コードの変更内容を、コサイン類似度で比較できます。

自然言語からのコード検索、コードからコードの検索、コード編集の検索に使えます。

項目内容
ベースモデル`Shuu12121/NightJar-large`
モデルの種類Sentence Transformer(Transformer + [CLS]プーリング)
パラメータ数約3.46億
出力次元1024
最大系列長1,024トークン
類似度コサイン類似度
タスクプレフィックスなし

評価結果(MTEB、nDCG@10)

MTEB 2.5.1で計算しました。比較対象のモデルも、すべて同じ手順で追加学習しています(詳細は「学習」を参照)。

モデルCodeSearchNetRetrieval(平均)CodeEditSearchRetrieval(平均)
NightJar-large-CodeSearch-Embedding0.91160.7822
NightJar-CodeSearch-Embedding0.90260.7382
NightOwl(同じ手順で追加学習)0.89970.7459

NightJar-large-CodeSearch-Embeddingは、CodeSearchNetRetrievalの6言語、CodeEditSearchRetrievalの13言語のすべてで、最も高いスコアでした。言語別のスコアは英語版の表を参照してください。

追加学習のデータは、NightOwl-CodeEmbeddingと同じ重複除去済みのデータセットです。学習前に、コード検索データとCodeSearchNetのテストデータの重複、commitpackft由来のコード編集データとCodeEditSearchRetrievalの評価データの重複を取り除いています。MTEBではCodeEditSearchRetrievalの評価データがtrainという名前になっていますが、この評価データは追加学習には使っていません。

使い方

英語版の「Usage」にサンプルコードがあります。

  • —最大長: 学習時の最大長は1,024トークンです。それより長い入力は切り捨てられます。
  • —双方向: 質問→コードと、コード→質問の両方向で学習しているので、コードに合う説明文を探す用途にも使えます。

学習

Shuu12121/NightJar-largeから、1エポック(3,978ステップ、4,073,472件)で追加学習しました。

データ: 各サンプルは、アンカー1件、正例1件、ハードネガティブ15件です。次の3つのデータセットを 8 : 8 : 1 の比率でサンプリングしました。

  • —owl_code_search_hard_negative_datasets_V2_kd(docstringからコードの検索、8言語)
  • —owl_code_search_hard_negative_datasets_additional_languages_V2_kd(同じく追加の8言語)
  • —codeedit_hard_negative_datasets_kd(コード編集の検索、11言語)

損失関数: Cached MultipleNegativesRankingLoss(scale 100、mini batch size 64、batch_size 1024)質問→文書と文書→質問の双方向で学習しています。

ハイパーパラメータ:

項目値
バッチサイズ1,024(勾配蓄積なし)
学習率6e-5(コサイン減衰、ウォームアップなし)
Weight decay0.01
最適化手法adamw_torch_fused
エポック数1
精度bf16
Gradient checkpointingあり
シード42

学習中の評価: 学習中は、Sentence TransformersのInformationRetrievalEvaluatorでCodeSearchNet(6言語、validation splitにおける性能)を監視しました。学習終了時のnDCG@10は、Go 0.9341、Java 0.8569、JavaScript 0.8420、PHP 0.7979、Python 0.9449、Ruby 0.9101で、平均は0.8810です(そのほかの指標は英語版の表を参照)。この数値はMTEBとは別の評価セットで計算したものなので、MTEBのスコアとは比較できません。

監視していたnDCG@10は、最初の数百ステップで急に上がり、約2,000ステップ(約0.5エポック)以降は、どの言語も最終値の±0.01以内で推移しました。

制限事項

  • —検索性能の評価は、CodeSearchNetRetrievalとCodeEditSearchRetrievalのみです。
  • —追加学習データの量や質などが言語ごとに違うため、性能にも差があります。追加学習データに含まれない言語では性能が下がる場合があります。
  • —1,024トークンを超える入力は切り捨てられます。
  • —検索で得られたコードには、バグ、古いAPI、安全でない実装が含まれる可能性があります。利用前に確認してください。

ライセンス

モデルの重みはApache 2.0で公開しています。学習データセットには、それぞれ独自のライセンスと利用条件があります。用途によっては、各データセットの条件も確認してください。