thealper2/codet5p-sql2text
SQL-to-Text (Salesforce/codet5p-220m)
Salesforce/codet5p-220m fine-tuned to explain a SQL query in plain English.
The direction is SQL -> natural language: the model takes a query (and, optionally, the DDL of the tables it touches) and returns a sentence describing what that query does. It does not generate SQL from a question.
Prompt format
Inputs follow one fixed template; training, evaluation and inference all build it with the same function, so they cannot drift apart. The schema block is dropped when no DDL is supplied, and when it is supplied only CREATE TABLE ... statements are kept.
Explain the following SQL query.
Schema:
<CREATE TABLE statements>
SQL:
<the query>Usage
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_id = "thealper2/codet5p-sql2text"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
schema = "CREATE TABLE employees (id INT, name TEXT, salary INT, dept_id INT);"
sql = "SELECT dept_id, AVG(salary) FROM employees GROUP BY dept_id;"
prompt = f"Explain the following SQL query.\n\nSchema:\n{schema}\n\nSQL:\n{sql}"
inputs = tokenizer(
prompt,
return_tensors="pt",
truncation=True,
max_length=256,
)
outputs = model.generate(
**inputs,
num_beams=4,
max_new_tokens=128,
min_new_tokens=5,
early_stopping=True,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Training data
`gretelai/synthetic_text_to_sql`, mapping sql + sql_context to sql_explanation.
Preprocessing drops rows that are too short to be a real explanation, removes exact duplicates and duplicate inputs, and removes any training row whose input also appears in the official test split, so the reported test scores are not inflated by leakage. The validation split is 3% of the cleaned train split (seed 42).
Sequence lengths were chosen from the measured token-length distribution: source 256 tokens, target 128 tokens.
Training procedure
Trained on a single NVIDIA GeForce RTX 5060 Ti (15.9 GB) with torch 2.11.0+cu128, bf16 mixed precision.
Wall-clock training time: 84 minutes.
Evaluation
Scored by evaluate.py on the full splits with beam search (num_beams=4).
Generation quality
SQL-aware faithfulness
Recall metrics ask whether the explanation mentions what the query actually does; the rate metrics are error rates, where lower is better -- they measure claims the query does not support.
Limitations
- Trained on synthetic queries and synthetic explanations, so the phrasing reflects that generator's style rather than how a particular team documents its own queries.
- Explanations are grounded in the query text, not in the data: the model cannot know what a column means beyond its name.
- Condition coverage is the weakest area -- long
WHEREclauses lose some columns and literals -- so an explanation may describe a filter less precisely than the query applies it. Do not rely on it as an audit of what a query returns. - Inputs are truncated past the configured source length, so very large schemas are only partially visible to the model.
- English only.
Reproduction
make preprocess
make train
make evaluateBase model: `Salesforce/codet5p-220m`.
