Team Ai
Modelpublic

dat-ai/bge-base-for_text2sql

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes78downloads
Model Card

BGE base SQL Matryoshka

This is a sentence-transformers model finetuned from BAAI/bge-base-en-v1.5 on the json dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: BAAI/bge-base-en-v1.5 <!-- at revision a5beb1e3e68b9ab74eb54cfd186867f64f240e1a -->
  • —Maximum Sequence Length: 512 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —json
  • —Language: en
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': True}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("dat-ai/bge-base-for_text2sql")
# Run inference
sentences = [
    '\n  Given  the Column informations, generate an SQL query for the following question:\n  Column: Nomination | Actors Name | Film Name | Director | Country\n  Question: What was the film Falling up nominated for?\n  SQL Query: SELECT Nomination FROM table WHERE Film Name = Falling Up\n  ',
    'What was the film Falling up nominated for?',
    'Who wrote an episode watched by 19.01 million US viewers?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
Metricdim_768dim_512dim_256dim_128dim_64
cosine_accuracy@10.46760.46780.46750.46770.4678
cosine_accuracy@30.46970.46970.46970.46960.4696
cosine_accuracy@50.46970.46970.46970.46980.4696
cosine_accuracy@100.46970.46970.46980.46980.4697
cosine_precision@10.46760.46780.46750.46770.4678
cosine_precision@30.15660.15660.15660.15650.1565
cosine_precision@50.09390.09390.09390.0940.0939
cosine_precision@100.0470.0470.0470.0470.047
cosine_recall@10.46760.46780.46750.46770.4678
cosine_recall@30.46970.46970.46970.46960.4696
cosine_recall@50.46970.46970.46970.46980.4696
cosine_recall@100.46970.46970.46980.46980.4697
cosine_ndcg@100.46890.4690.46890.46890.469
cosine_mrr@100.46860.46870.46860.46870.4687
cosine_map@1000.46860.46870.46860.46870.4687

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

json
  • —Dataset: json
  • —Size: 56,355 training samples
  • —Columns: <code>context</code> and <code>question</code>
  • —Approximate statistics based on the first 1000 samples: | | context | question | |:--------|:------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 45 tokens</li><li>mean: 72.61 tokens</li><li>max: 196 tokens</li></ul> | <ul><li>min: 7 tokens</li><li>mean: 15.41 tokens</li><li>max: 36 tokens</li></ul> |
  • —Samples: | context | question | |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------| | <code><br> Given the Column informations, generate an SQL query for the following question:<br> Column: State/territory | Text/background colour | Format | Current slogan | Current series | Notes<br> Question: Tell me what the notes are for South Australia <br> SQL Query: SELECT Notes FROM table WHERE Current slogan = SOUTH AUSTRALIA<br> </code> | <code>Tell me what the notes are for South Australia </code> | | <code><br> Given the Column informations, generate an SQL query for the following question:<br> Column: State/territory | Text/background colour | Format | Current slogan | Current series | Notes<br> Question: What is the current series where the new series began in June 2011?<br> SQL Query: SELECT Current series FROM table WHERE Notes = New series began in June 2011<br> </code> | <code>What is the current series where the new series began in June 2011?</code> | | <code><br> Given the Column informations, generate an SQL query for the following question:<br> Column: State/territory | Text/background colour | Format | Current slogan | Current series | Notes<br> Question: What is the format for South Australia?<br> SQL Query: SELECT Format FROM table WHERE State/territory = South Australia<br> </code> | <code>What is the format for South Australia?</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512
      ],
      "matryoshka_weights": [
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 16
  • —gradient_accumulation_steps: 8
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —fp16: True
  • —load_best_model_at_end: True
  • —optim: adamwtorchfused
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 16
  • —per_device_eval_batch_size: 8
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 8
  • —eval_accumulation_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: False
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —dispatch_batches: None
  • —split_batches: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

<details><summary>Click to expand</summary>

EpochStepTraining Lossdim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.0227101.773-----
0.0454201.3231-----
0.0681300.713-----
0.0908400.286-----
0.1135500.1013-----
0.1362600.0635-----
0.1590700.0453-----
0.1817800.041-----
0.2044900.039-----
0.22711000.027-----
0.24981100.0193-----
0.27251200.0167-----
0.29521300.016-----
0.31791400.0197-----
0.34061500.0217-----
0.36331600.0162-----
0.38601700.012-----
0.40871800.013-----
0.43151900.0255-----
0.45422000.0229-----
0.47692100.0181-----
0.49962200.0195-----
0.52232300.0199-----
0.54502400.0144-----
0.56772500.0102-----
0.59042600.0101-----
0.61312700.0095-----
0.63582800.0173-----
0.65852900.01-----
0.68123000.0129-----
0.70393100.0177-----
0.72673200.0106-----
0.74943300.0146-----
0.77213400.0185-----
0.79483500.0203-----
0.81753600.0146-----
0.84023700.0072-----
0.86293800.0102-----
0.88563900.0075-----
0.90834000.0064-----
0.93104100.0163-----
0.95374200.0069-----
0.97644300.0072-----
0.99914400.01470.46880.46890.46880.46890.4689
1.02194500.0151-----
1.04464600.0135-----
1.06734700.0189-----
1.09004800.0121-----
1.11274900.0064-----
1.13545000.0111-----
1.15815100.0103-----
1.18085200.0144-----
1.20355300.0151-----
1.22625400.0062-----
1.24895500.0104-----
1.27165600.0046-----
1.29445700.0056-----
1.31715800.0073-----
1.33985900.007-----
1.36256000.0074-----
1.38526100.0057-----
1.40796200.0052-----
1.43066300.0114-----
1.45336400.0075-----
1.47606500.0116-----
1.49876600.0092-----
1.52146700.0137-----
1.54416800.0066-----
1.56686900.0042-----
1.58967000.0036-----
1.61237100.0039-----
1.63507200.0065-----
1.65777300.0051-----
1.68047400.0054-----
1.70317500.0086-----
1.72587600.0062-----
1.74857700.0071-----
1.77127800.0108-----
1.79397900.009-----
1.81668000.0075-----
1.83938100.0039-----
1.86208200.0047-----
1.88488300.0037-----
1.90758400.0037-----
1.93028500.0064-----
1.95298600.0047-----
1.97568700.0034-----
1.99838800.00610.46890.46890.46890.46900.4690
2.02108900.0096-----
2.04379000.0071-----
2.06649100.0101-----
2.08919200.0054-----
2.11189300.0039-----
2.13459400.0074-----
2.15739500.0044-----
2.18009600.0088-----
2.20279700.0096-----
2.22549800.0057-----
2.24819900.0063-----
2.270810000.0026-----
2.293510100.0032-----
2.316210200.0027-----
2.338910300.0041-----
2.361610400.0052-----
2.384310500.0035-----
2.407010600.0025-----
2.429710700.0059-----
2.452510800.0048-----
2.475210900.0064-----
2.497911000.0066-----
2.520611100.0078-----
2.543311200.0057-----
2.566011300.0026-----
2.588711400.0021-----
2.611411500.0021-----
2.634111600.0047-----
2.656811700.0034-----
2.679511800.0044-----
2.702211900.0058-----
2.725012000.0043-----
2.747712100.0056-----
2.770412200.0076-----
2.793112300.0063-----
2.815812400.0033-----
2.838512500.0025-----
2.861212600.0019-----
2.883912700.0052-----
2.906612800.0021-----
2.929312900.0041-----
2.952013000.0035-----
2.974713100.0044-----
2.997413200.0035-----
2.99971321-0.4690.4690.4690.4690.469
3.020213300.0062-----
3.042913400.0047-----
3.065613500.008-----
3.088313600.0033-----
3.111013700.0025-----
3.133713800.0069-----
3.156413900.0035-----
3.179114000.0085-----
3.201814100.007-----
3.224514200.007-----
3.247214300.0052-----
3.269914400.0019-----
3.292614500.0022-----
3.315414600.0019-----
3.338114700.0028-----
3.360814800.0042-----
3.383514900.0023-----
3.406215000.0024-----
3.428915100.0036-----
3.451615200.0038-----
3.474315300.0063-----
3.497015400.0044-----
3.519715500.0064-----
3.542415600.0053-----
3.565115700.0019-----
3.587915800.0019-----
3.610615900.0017-----
3.633316000.004-----
3.656016100.0026-----
3.678716200.0031-----
3.701416300.0043-----
3.724116400.0032-----
3.746816500.0041-----
3.769516600.0069-----
3.792216700.0063-----
3.814916800.0038-----
3.837616900.0024-----
3.860317000.0018-----
3.883117100.0034-----
3.905817200.0016-----
3.928517300.0026-----
3.951217400.0037-----
3.973917500.0024-----
3.996617600.00270.46890.46900.46890.46890.4690
  • —The bold row denotes the saved checkpoint. </details>

Framework Versions

  • —Python: 3.10.14
  • —Sentence Transformers: 3.3.0
  • —Transformers: 4.41.2
  • —PyTorch: 2.1.2+cu121
  • —Accelerate: 0.34.2
  • —Datasets: 2.19.1
  • —Tokenizers: 0.19.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->