Team Ai
Datasetpublic

acomagu/dbbench-mysql-synth

DBBench-Style MySQL Synthetic SFT Dataset (English) This directory contains a DBBench-style (MySQL / English) synthetic dataset for SFT. It is designed for: SFT training in a DBBench-like interaction format Auditing data quality (SQL / answers / difficulty / distribution) Reproducing and debugging queries on MySQL Future DAgger / hard example mining This dataset is synthetic and not the official DBBench data. It is meant to teach DBBench-style skills (SQL generation, result… See the full description on the dataset page: https://huggingface.co/datasets/acomagu/dbbench-mysql-synth.

sourceHugging Faceotherupdated 7mo agoView on Hugging Face
0likes33downloads
Dataset Card

DBBench-Style MySQL Synthetic SFT Dataset (English)

This directory contains a DBBench-style (MySQL / English) synthetic dataset for SFT. It is designed for:

  • —SFT training in a DBBench-like interaction format
  • —Auditing data quality (SQL / answers / difficulty / distribution)
  • —Reproducing and debugging queries on MySQL
  • —Future DAgger / hard example mining

This dataset is synthetic and not the official DBBench data. It is meant to teach DBBench-style skills (SQL generation, result reading, format adherence, and recovery behavior).

Directory Structure

  • —catalog/
  • —manifest.json
  • —data.jsonl
  • —messages_val.jsonl
  • —mysql_seed/
  • —tasks_all.jsonl
  • —tasks_train.jsonl
  • —task_summary.csv
  • —tasks_val.jsonl

What to Use for What

For SFT training (main inputs)

  • —data.jsonl (train)
  • —messages_val.jsonl (validation)

For audit / analysis (structured metadata)

  • —tasks_train.jsonl
  • —tasks_val.jsonl
  • —tasks_all.jsonl
  • —task_summary.csv

For reproduction / debugging (reloadable into MySQL)

  • —mysql_seed/

For schema inspection / quality checks

  • —catalog/

For generation conditions / reproducibility

  • —manifest.json

1. data.jsonl / messages_val.jsonl

Role

These are the primary SFT inputs. Each line is a JSON object with a messages array.

Key characteristics

  • —DBBench-style conversation format
  • —English questions / reasoning / final answers
  • —Action: Operation (SQL execution)
  • —User-side messages include raw MySQL-like responses
  • —Action: Answer / Final Answer: [...]
  • —Some samples include real error → fix → re-run trajectories

Notes

Some notebooks will ignore messages_val.jsonl and do their own split. In that case, a common workflow is:

  1. 1.Upload data.jsonl to HF
  2. 2.Let the notebook split with SFT_VAL_RATIO

2. tasks_train.jsonl / tasks_val.jsonl / tasks_all.jsonl

Role

Structured metadata for auditing and analysis. Easier to work with than messages_*.

Typical fields include:

  • —task_id, db_id, domain
  • —question_en, gold_sql
  • —difficulty, tags, sql_tables
  • —execution_validation.final_answer

Common uses

  • —SQL validity checks
  • —Difficulty / domain distribution
  • —Hard example mining (e.g., subquery, having, quoted_identifier)
  • —Quick evaluation after SFT

3. task_summary.csv

A human-readable summary for quick inspection in Excel / Sheets / pandas:

  • —domain balance
  • —difficulty balance
  • —error recovery ratio
  • —rowcount distribution
  • —surface-level duplication checks

4. catalog/

Schema metadata per db_id (table/column info and row counts), useful for:

  • —Understanding schemas
  • —Auditing column types
  • —Prompt design and future data generation

5. mysql_seed/

MySQL re-import SQL (DDL + INSERT) for each db_id:

  • —Reproduce specific tasks
  • —Debug model SQL with real execution
  • —Rebuild datasets for DAgger

Example:

bash
mysql -u root -p dbbench_syn_ecommerce_0001 < mysql_seed/ecommerce_0001.sql

6. manifest.json

Generation conditions for reproducibility and auditing:

  • —generation time
  • —dialect
  • —n_dbs_generated, n_tasks_total
  • —train/val sizes
  • —domains
  • —error recovery ratio
  • —db_id splits

Minimal Share Set

If you only need SFT training:

  • —data.jsonl
  • —messages_val.jsonl
  • —manifest.json

If you also want audit:

  • —plus tasks_train.jsonl, tasks_val.jsonl, task_summary.csv

If you need full reproduction:

  • —plus catalog/ and mysql_seed/

Quick sanity checks

bash
wc -l data.jsonl messages_val.jsonl
head -n 2 data.jsonl
bash
wc -l tasks_all.jsonl tasks_train.jsonl tasks_val.jsonl
head -n 2 tasks_train.jsonl
bash
head -n 20 task_summary.csv

Notes and Limitations

  • —Synthetic dataset; not official DBBench data.
  • —Train/val split is done by db_id to reduce leakage across databases.
  • —messages_val.jsonl may be unused if your training pipeline re-splits train data.

If you find gaps (e.g., too few hard cases, null/empty results, edge cases), consider regenerating with adjusted parameters.