Team Ai
Modelpublic

thealper2/codet5p-sql2text

sourceHugging Facebsd-3-clauseupdated 15d agoView on Hugging Face
0likes298downloads
Model Card

SQL-to-Text (Salesforce/codet5p-220m)

Salesforce/codet5p-220m fine-tuned to explain a SQL query in plain English.

The direction is SQL -> natural language: the model takes a query (and, optionally, the DDL of the tables it touches) and returns a sentence describing what that query does. It does not generate SQL from a question.

Prompt format

Inputs follow one fixed template; training, evaluation and inference all build it with the same function, so they cannot drift apart. The schema block is dropped when no DDL is supplied, and when it is supplied only CREATE TABLE ... statements are kept.

Explain the following SQL query.

Schema:
<CREATE TABLE statements>

SQL:
<the query>

Usage

python
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

model_id = "thealper2/codet5p-sql2text"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

schema = "CREATE TABLE employees (id INT, name TEXT, salary INT, dept_id INT);"
sql = "SELECT dept_id, AVG(salary) FROM employees GROUP BY dept_id;"
prompt = f"Explain the following SQL query.\n\nSchema:\n{schema}\n\nSQL:\n{sql}"

inputs = tokenizer(
    prompt,
    return_tensors="pt",
    truncation=True,
    max_length=256,
)
outputs = model.generate(
    **inputs,
    num_beams=4,
    max_new_tokens=128,
    min_new_tokens=5,
    early_stopping=True,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training data

`gretelai/synthetic_text_to_sql`, mapping sql + sql_context to sql_explanation.

Preprocessing drops rows that are too short to be a real explanation, removes exact duplicates and duplicate inputs, and removes any training row whose input also appears in the official test split, so the reported test scores are not inflated by leakage. The validation split is 3% of the cleaned train split (seed 42).

Sequence lengths were chosen from the measured token-length distribution: source 256 tokens, target 128 tokens.

Training procedure

Hyper-parameterValue
Epochs3.00
Learning rate0.0001
LR schedulelinear
Warmup ratio0.0500
Weight decay0.0100
Optimiseradamw_torch
Per-device train batch size16
Gradient accumulation4
Max gradient norm1.00
Model selectioneval_rougeL
Seed42
Effective batch size64

Trained on a single NVIDIA GeForce RTX 5060 Ti (15.9 GB) with torch 2.11.0+cu128, bf16 mixed precision.

Wall-clock training time: 84 minutes.

Evaluation

Scored by evaluate.py on the full splits with beam search (num_beams=4).

Generation quality

MetricValidationTest_meta
Examples29895850-
BLEU33.6533.10-
ROUGE-167.1566.86-
ROUGE-245.7645.11-
ROUGE-L57.5857.06-
Mean generated length36.3735.69-
Loss0.58290.5923-

SQL-aware faithfulness

Recall metrics ask whether the explanation mentions what the query actually does; the rate metrics are error rates, where lower is better -- they measure claims the query does not support.

MetricValidationTest
Examples29895850
Operation recall98.5398.40
Aggregation recall98.7298.57
Join mention recall98.2198.74
Join table coverage98.3498.56
Condition column coverage82.9881.99
Condition value coverage86.1887.24
Operation over-claim rate5.475.29
Unsupported number rate3.983.18
Unsupported quoted-string rate3.203.18
Unsupported entity rate1.381.72

Limitations

  • —Trained on synthetic queries and synthetic explanations, so the phrasing reflects that generator's style rather than how a particular team documents its own queries.
  • —Explanations are grounded in the query text, not in the data: the model cannot know what a column means beyond its name.
  • —Condition coverage is the weakest area -- long WHERE clauses lose some columns and literals -- so an explanation may describe a filter less precisely than the query applies it. Do not rely on it as an audit of what a query returns.
  • —Inputs are truncated past the configured source length, so very large schemas are only partially visible to the model.
  • —English only.

Reproduction

bash
make preprocess
make train
make evaluate

Base model: `Salesforce/codet5p-220m`.