acomagu/dbbench-mysql-synth
DBBench-Style MySQL Synthetic SFT Dataset (English) This directory contains a DBBench-style (MySQL / English) synthetic dataset for SFT. It is designed for: SFT training in a DBBench-like interaction format Auditing data quality (SQL / answers / difficulty / distribution) Reproducing and debugging queries on MySQL Future DAgger / hard example mining This dataset is synthetic and not the official DBBench data. It is meant to teach DBBench-style skills (SQL generation, result… See the full description on the dataset page: https://huggingface.co/datasets/acomagu/dbbench-mysql-synth.
DBBench-Style MySQL Synthetic SFT Dataset (English)
This directory contains a DBBench-style (MySQL / English) synthetic dataset for SFT. It is designed for:
- SFT training in a DBBench-like interaction format
- Auditing data quality (SQL / answers / difficulty / distribution)
- Reproducing and debugging queries on MySQL
- Future DAgger / hard example mining
This dataset is synthetic and not the official DBBench data. It is meant to teach DBBench-style skills (SQL generation, result reading, format adherence, and recovery behavior).
Directory Structure
catalog/manifest.jsondata.jsonlmessages_val.jsonlmysql_seed/tasks_all.jsonltasks_train.jsonltask_summary.csvtasks_val.jsonl
What to Use for What
For SFT training (main inputs)
data.jsonl(train)messages_val.jsonl(validation)
For audit / analysis (structured metadata)
tasks_train.jsonltasks_val.jsonltasks_all.jsonltask_summary.csv
For reproduction / debugging (reloadable into MySQL)
mysql_seed/
For schema inspection / quality checks
catalog/
For generation conditions / reproducibility
manifest.json
1. data.jsonl / messages_val.jsonl
Role
These are the primary SFT inputs. Each line is a JSON object with a messages array.
Key characteristics
- DBBench-style conversation format
- English questions / reasoning / final answers
Action: Operation(SQL execution)- User-side messages include raw MySQL-like responses
Action: Answer/Final Answer: [...]- Some samples include real error → fix → re-run trajectories
Notes
Some notebooks will ignore messages_val.jsonl and do their own split. In that case, a common workflow is:
- Upload
data.jsonlto HF - Let the notebook split with
SFT_VAL_RATIO
2. tasks_train.jsonl / tasks_val.jsonl / tasks_all.jsonl
Role
Structured metadata for auditing and analysis. Easier to work with than messages_*.
Typical fields include:
task_id,db_id,domainquestion_en,gold_sqldifficulty,tags,sql_tablesexecution_validation.final_answer
Common uses
- SQL validity checks
- Difficulty / domain distribution
- Hard example mining (e.g.,
subquery,having,quoted_identifier) - Quick evaluation after SFT
3. task_summary.csv
A human-readable summary for quick inspection in Excel / Sheets / pandas:
- domain balance
- difficulty balance
- error recovery ratio
- rowcount distribution
- surface-level duplication checks
4. catalog/
Schema metadata per db_id (table/column info and row counts), useful for:
- Understanding schemas
- Auditing column types
- Prompt design and future data generation
5. mysql_seed/
MySQL re-import SQL (DDL + INSERT) for each db_id:
- Reproduce specific tasks
- Debug model SQL with real execution
- Rebuild datasets for DAgger
Example:
mysql -u root -p dbbench_syn_ecommerce_0001 < mysql_seed/ecommerce_0001.sql6. manifest.json
Generation conditions for reproducibility and auditing:
- generation time
- dialect
n_dbs_generated,n_tasks_total- train/val sizes
- domains
- error recovery ratio
- db_id splits
Minimal Share Set
If you only need SFT training:
data.jsonlmessages_val.jsonlmanifest.json
If you also want audit:
- plus
tasks_train.jsonl,tasks_val.jsonl,task_summary.csv
If you need full reproduction:
- plus
catalog/andmysql_seed/
Quick sanity checks
wc -l data.jsonl messages_val.jsonl
head -n 2 data.jsonlwc -l tasks_all.jsonl tasks_train.jsonl tasks_val.jsonl
head -n 2 tasks_train.jsonlhead -n 20 task_summary.csvNotes and Limitations
- Synthetic dataset; not official DBBench data.
- Train/val split is done by
db_idto reduce leakage across databases. messages_val.jsonlmay be unused if your training pipeline re-splits train data.
If you find gaps (e.g., too few hard cases, null/empty results, edge cases), consider regenerating with adjusted parameters.
