Team Ai
Modelpublic

KayaTechAI/Qwen3-0.6B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-Generalization-Baseline

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes112downloads
Model Card

Qwen3-Telecom-Retrieval-Embedding

This is a sentence-transformers model finetuned from Qwen/Qwen3-Embedding-0.6B on the telecom-technical-documents-retrieval-embedding-dataset dataset. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 32768, 'do_lower_case': False, 'architecture': 'Qwen3Model'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': True, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("KayaTechAI/Qwen3-0.6B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-Generalization-Baseline")
# Run inference
queries = [
    "What is the provisioning scope for the eMLPP service?",
]
documents = [
    'eMLPP is provisioned per subscriber.',
    'The main objective is to verify that the User Equipment (UE) tracks channel variations and selects the optimal transport format for frequency non-selective scheduling.',
    'SDP is used in SIP communications to describe the parameters and media capabilities of a session, such as audio/video codecs, transport protocols, and IP addresses, enabling participants to agree on the media types to be used.',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 1024] [3, 1024]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.6303, -0.0008, -0.0340]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 1024
  }
MetricValue
cosine_accuracy@10.7988
cosine_accuracy@30.912
cosine_accuracy@50.9404
cosine_accuracy@100.9636
cosine_precision@10.7988
cosine_precision@30.304
cosine_precision@50.1881
cosine_precision@100.0964
cosine_recall@10.7988
cosine_recall@30.912
cosine_recall@50.9404
cosine_recall@100.9636
cosine_ndcg@100.886
cosine_mrr@100.8606
cosine_map@1000.8621
Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.7996
cosine_accuracy@30.9148
cosine_accuracy@50.9408
cosine_accuracy@100.9624
cosine_precision@10.7996
cosine_precision@30.3049
cosine_precision@50.1882
cosine_precision@100.0962
cosine_recall@10.7996
cosine_recall@30.9148
cosine_recall@50.9408
cosine_recall@100.9624
cosine_ndcg@100.8859
cosine_mrr@100.8608
cosine_map@1000.8625
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.7968
cosine_accuracy@30.9128
cosine_accuracy@50.9388
cosine_accuracy@100.962
cosine_precision@10.7968
cosine_precision@30.3043
cosine_precision@50.1878
cosine_precision@100.0962
cosine_recall@10.7968
cosine_recall@30.9128
cosine_recall@50.9388
cosine_recall@100.962
cosine_ndcg@100.8844
cosine_mrr@100.8589
cosine_map@1000.8606
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.7804
cosine_accuracy@30.912
cosine_accuracy@50.9316
cosine_accuracy@100.9584
cosine_precision@10.7804
cosine_precision@30.304
cosine_precision@50.1863
cosine_precision@100.0958
cosine_recall@10.7804
cosine_recall@30.912
cosine_recall@50.9316
cosine_recall@100.9584
cosine_ndcg@100.8753
cosine_mrr@100.848
cosine_map@1000.8496
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.7696
cosine_accuracy@30.898
cosine_accuracy@50.9268
cosine_accuracy@100.9524
cosine_precision@10.7696
cosine_precision@30.2993
cosine_precision@50.1854
cosine_precision@100.0952
cosine_recall@10.7696
cosine_recall@30.898
cosine_recall@50.9268
cosine_recall@100.9524
cosine_ndcg@100.8663
cosine_mrr@100.8381
cosine_map@1000.8399
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.75
cosine_accuracy@30.8816
cosine_accuracy@50.9124
cosine_accuracy@100.9456
cosine_precision@10.75
cosine_precision@30.2939
cosine_precision@50.1825
cosine_precision@100.0946
cosine_recall@10.75
cosine_recall@30.8816
cosine_recall@50.9124
cosine_recall@100.9456
cosine_ndcg@100.8522
cosine_mrr@100.8218
cosine_map@1000.8236

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

telecom-technical-documents-retrieval-embedding-dataset
  • —Dataset: telecom-technical-documents-retrieval-embedding-dataset at 3ebf34a
  • —Size: 127,731 training samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:----------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 7 tokens</li><li>mean: 18.79 tokens</li><li>max: 68 tokens</li></ul> | <ul><li>min: 4 tokens</li><li>mean: 26.09 tokens</li><li>max: 77 tokens</li></ul> |
  • —Samples: | anchor | positive | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>What is the estimated Transmit power considered sufficient for achieving 95% Downlink coverage with a single Base Station?</code> | <code>Approximately 14 dBm Transmit power is considered sufficient.</code> | | <code>What is the primary goal of the Nominal Accuracy requirement?</code> | <code>The primary goal of the Nominal Accuracy requirement is to ensure good accuracy when signal conditions are ideal.</code> | | <code>What happens on the mobile station side if contention resolution fails because the G-RNTI value in the network's acknowledgement message differs from what the mobile station sent?</code> | <code>If the mobile station receives a PACKET UPLINK ACK/NACK message with a G-RNTI value different from the one it included in its first RLC data blocks, it signifies a contention resolution failure, and the mobile station will not transmit a PACKET CONTROL ACKNOWLEDGEMENT.</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          1024,
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 32
  • —gradient_accumulation_steps: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —bf16: True
  • —tf32: True
  • —load_best_model_at_end: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 32
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 16
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: True
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Lossdim_1024_cosine_ndcg@10dim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.0401101.5256------
0.0802200.8247------
0.1202300.4102------
0.1603400.27------
0.2004500.2182------
0.2405600.1998------
0.2806700.2017------
0.3206800.1672------
0.3607900.2029------
0.40081000.1609------
0.44091100.1565------
0.48101200.1476------
0.52101300.1278------
0.56111400.1669------
0.60121500.1642------
0.64131600.1307------
0.68141700.1487------
0.72141800.1329------
0.76151900.13------
0.80162000.1393------
0.84172100.1344------
0.88182200.1184------
0.92182300.1147------
0.96192400.1283------
1.02500.12280.86930.86830.86340.85350.84300.8082
1.04012600.0613------
1.08022700.0559------
1.12022800.0704------
1.16032900.0578------
1.20043000.0588------
1.24053100.079------
1.28063200.0602------
1.32063300.0553------
1.36073400.0663------
1.40083500.0513------
1.44093600.0615------
1.48103700.0462------
1.52103800.0674------
1.56113900.0558------
1.60124000.0562------
1.64134100.0688------
1.68144200.0905------
1.72144300.0463------
1.76154400.0581------
1.80164500.0586------
1.84174600.0712------
1.88184700.041------
1.92184800.0578------
1.96194900.063------
2.05000.05050.87710.87800.87640.86900.85870.8353
2.04015100.032------
2.08025200.0239------
2.12025300.029------
2.16035400.0236------
2.20045500.0381------
2.24055600.028------
2.28065700.0366------
2.32065800.0372------
2.36075900.0306------
2.40086000.0294------
2.44096100.0269------
2.48106200.0411------
2.52106300.0251------
2.56116400.0299------
2.60126500.0275------
2.64136600.0267------
2.68146700.0304------
2.72146800.0246------
2.76156900.025------
2.80167000.037------
2.84177100.0393------
2.88187200.0405------
2.92187300.0279------
2.96197400.0243------
3.07500.02840.88700.88580.88270.87450.86480.8499
3.04017600.0166------
3.08027700.024------
3.12027800.0302------
3.16037900.0263------
3.20048000.0172------
3.24058100.023------
3.28068200.0313------
3.32068300.0253------
3.36078400.0189------
3.40088500.0177------
3.44098600.0187------
3.48108700.0142------
3.52108800.0281------
3.56118900.0253------
3.60129000.0184------
3.64139100.0217------
3.68149200.027------
3.72149300.0192------
3.76159400.0183------
3.80169500.0242------
3.84179600.0223------
3.88189700.0161------
3.92189800.0219------
3.96199900.0236------
4.010000.02780.8860.88590.88440.87530.86630.8522
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.12.12
  • —Sentence Transformers: 5.2.3
  • —Transformers: 4.55.4
  • —PyTorch: 2.10.0+cu128
  • —Accelerate: 1.12.0
  • —Datasets: 3.6.0
  • —Tokenizers: 0.21.4

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->