MaXoN654/RUSQL-0.8B-Text2SQL
RUSQL-0.8B-Text2SQL
Compact Russian text-to-SQL model: a full-parameter SFT of techwithsergiu/Qwen3.5-text-0.8B — a text-only slice of Qwen/Qwen3.5-0.8B with the vision tower removed (0.77B actual parameters) — trained to answer Russian natural-language questions over a database schema with step-by-step reasoning that ends in a final SQLite query (OmniSQL-style CoT).
Performance Evaluation
Execution accuracy (predicted SQL executed against SQLite, result-set comparison) on the held-out eval split — the same 2,729 questions in every row (EN vs RU are the same items, English question vs its translation), greedy decoding:
Fine-tuning lifts execution accuracy from 13.9% to 58.4% — about 4.2× the base model, and above its English-question ceiling (16.0%).
Breakdown by SQL complexity (RU questions):
All three rows are scored on the exact same 2,729 items (the QE-filtered held-out split), so the numbers are directly comparable. Greedy decoding. The base model also fails to produce a parseable SQL block on ~12% of items (331/2,729 for RU, 322 for EN); after fine-tuning this drops to 13.
Dataset Overview
Training data is derived from SynSQL-2.5M (OmniSQL, arXiv:2503.02240) through a fully local, streaming pipeline:
Only the question is translated to Russian; schema (DDL), external knowledge and the gold SQL stay in English — matching the real-world setting where databases are English-named but users ask in Russian.
- Training set: ~444k filtered examples (+2,729 held-out eval); filter drop rate ~13.8%
Instruction Prompt
The model is trained (and must be used) with this exact chat format:
System:
You are a text-to-SQL assistant. Given a database schema and a question, reason step by step and finish with the final SQLite query in a ```sql code block.User:
Database schema:
{DDL}
External knowledge:
{optional, may be omitted}
Question: {вопрос на русском}Assistant: free-form chain-of-thought ending with the final query in a `sql ... ` block. Qwen thinking mode is disabled (enable_thinking=False) — reasoning is plain response text, OmniSQL style.
Training Configuration
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import re, torch
model_id = "MaXoN654/RUSQL-0.8B-Text2SQL"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")
# schema = "\n\n".join of CREATE TABLE statements, SynSQL style
# (quoted identifiers, inline /* ... */ column comments)
schema = """CREATE TABLE "employees" (
"employee_id" INTEGER /* Unique identifier for each employee */,
"name" TEXT /* Full name of the employee */,
"salary" REAL /* Annual salary in USD */,
"department_id" INTEGER /* Reference to the department */,
PRIMARY KEY ("employee_id"),
CONSTRAINT fk_employees_department_id FOREIGN KEY ("department_id") REFERENCES departments ("department_id")
)
CREATE TABLE "departments" (
"department_id" INTEGER /* Unique identifier for each department */,
"department_name" TEXT /* Name of the department */,
PRIMARY KEY ("department_id")
)"""
question = "Покажи трёх сотрудников с самой высокой зарплатой в отделе продаж"
external_knowledge = None # optional hint text; omitted from the prompt when empty
def build_user(schema, question, external_knowledge=None):
parts = [f"Database schema:\n{schema}"]
if external_knowledge and external_knowledge.strip():
parts.append(f"External knowledge:\n{external_knowledge.strip()}")
parts.append(f"Question: {question}")
return "\n\n".join(parts)
messages = [
{"role": "system", "content": "You are a text-to-SQL assistant. Given a database schema and a question, reason step by step and finish with the final SQLite query in a ```sql code block."},
{"role": "user", "content": build_user(schema, question, external_knowledge)},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
enable_thinking=False, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=1024, temperature=0.0, do_sample=False)
text = tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True)
sql = re.findall(r"```sql\s*(.*?)```", text, re.S | re.I)[-1].strip()
print(sql)Limitations
- SELECT-only — SynSQL-2.5M contains no DML/DDL, so the model was never trained on INSERT/UPDATE/DELETE. Asked to modify data, it does not refuse — it silently reformulates the request into a related SELECT. Do not use it to generate data-modifying queries.
- SQLite dialect only — queries may not be valid PostgreSQL/MySQL without adaptation.
- Schema and gold SQL are English; questions in other languages than Russian/English are untested.
- 0.8B parameters: complex multi-join / nested queries remain challenging; verify results before use.
- Training questions are machine-translated — residual translation artifacts are possible despite QE filtering.
Pipeline
The full data pipeline (sampling → translation → QE filtering → SFT → execution-accuracy eval) is implemented in the rusql project and runs entirely on a single consumer GPU.
