datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.Text-to-sql-v1enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,087 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a
self-correction loop — and the benchmark is designed so that number cannot be
inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.bird-sqlite-sft-train
BIRD SQLite Text-to-SQL SFT Dataset
Supervised fine-tuning data for a SQLite-dialect text-to-SQL specialist model,
built from the BIRD benchmark train split.
Contents
7,483 train + 408 val examples spanning 69 distinct database schemas
(movie_platform, chicago_crime, hockey, mondial_geo, works_cycles, and 64
others), split by a stratified per-database 95/5 hold-out (sft_sqlite_ train.jsonl / sft_sqlite_val.jsonl) with zero exact overlap between them.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/hiimivantang/bird-sqlite-sft-train.Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.Effi-SQL
Effi-SQL
Update 2026-06-12
We release Effi-SQL, a dataset suite for SQL efficiency optimization.
This collection includes:
Effi-SQL Benchmark: a benchmark for evaluating SQL efficiency optimization methods.
Diff-SQL Training Dataset: training data used by Diff-SQL, including data for the Patch Generator and Constraint Aligner.
Dataset Fields
Effi-SQL Benchmark
id: A unique identifier for each benchmark instance.
db: The database… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/Effi-SQL.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.SQLFlow
Text2SQL-Flow Dataset Repository
This repository contains the SQLFlow dataset.
The SQLFlow dataset is a large-scale, high-quality collection of semantically valid and structurally diverse Text-to-SQL examples, generated using a comprehensive SQL-aware data augmentation framework.
For more details, please visit the GitHub repository:🔗 https://github.com/TechNomad-ds/Text2SQL-Flow
mirror-sql
MIRROR-SQL
Provenance-Controlled Database Environments for Text-to-SQL Agents.
13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs.
MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus.
Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be
redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources.
MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.ap-sql-peft
ap-sql-peft
A chat-format text-to-SQL dataset for Accounts Payable analytics on Oracle. The schema is modelled on Oracle Fusion
AP, Payments and Supplier tables, and the data behind it is synthetic. The dataset was used to train the LoRA adapter
samrat-kar/ap-sql-v1.
Every gold SQL query was run against the demo database when the dataset was built; meta.result_rows records how many rows it returned.
Format
One JSON object per line:
{"messages": [{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/samrat-kar/ap-sql-peft.database-sql-instruction-dataset
Database & SQL Instruction Dataset
High-quality instruction-response pairs covering PostgreSQL, advanced queries, indexing strategies, and database optimization.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Database topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/database-sql-instruction-dataset.sql-parsedsql-query-generation-sft-100k
SQL Query Generation SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality SQL query generation from natural language requests. Each example includes a realistic database schema, a natural language query request, a correct SQL query, and a clear explanation of how the query works — across 6 SQL dialects and 15+ complexity levels.
Motivation
Text-to-SQL is one of the highest-value NLP applications in enterprise settings. Common model failures… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sql-query-generation-sft-100k.Text-to-sql-v1sql-query-engine-synthetic
SQL Query Engine — Synthetic Benchmark
A gold-standard NL-to-SQL benchmark containing 75 natural language questions across 3 PostgreSQL databases (e-commerce, university, hospital), each with verified gold SQL queries and expected results. Designed to evaluate text-to-SQL systems with a focus on measuring the impact of iterative self-healing (query repair) loops.
Paper
SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation
Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/codeadeel/sql-query-engine-synthetic.sql-create-context-id
Overview
This dataset is a fork from sql-create-context
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/detakarang/sql-create-context-id.pt-br-agentic-text-to-sql-distilled-trajectories
PT-BR Agentic Text-to-SQL Distilled Trajectories
This dataset contains message-only distilled trajectories for training tool-using Text-to-SQL agents in Brazilian Portuguese. The trajectories were selected from LLM-judged correct conversations and preserve the agent protocol used in the released code.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/pt-br-agentic-text-to-sql-distilled-trajectories.Cosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
🖥️ Demo Interface: Discord
Discord: https://discord.gg/Xe9tHFCS9h
**Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.sql_bi__b_db
Набор данных: timbossm/sql_bi__b_db
Этот набор данных содержит SQL-запросы и соответствующий им контекст баз данных, разработанный на основе методического пособия "Лабораторный практикум по языку SQL: практикум" (сост. Т. М. Босенко, Ю.В. Фролов. – М.: МГПУ, 2025. – 101 с.).
Набор данных включает 25 вариантов схем баз данных и предназначен для обучения и практики в написании SQL-запросов. Он охватывает следующие темы лабораторных работ:
ЛАБОРАТОРНАЯ РАБОТА №1. Изучение команд DDL.… See the full description on the dataset page: https://huggingface.co/datasets/timbossm/sql_bi__b_db.synthetic-text-to-sql-tangle
Synthetic Text-to-SQL — Tangle SFT demo subset
A small, fixed subset of gretelai/synthetic_text_to_sql
(Apache-2.0), prepared for an end-to-end supervised fine-tuning showcase running as a
Tangle pipeline.
train.jsonl — first 5,000 rows of the source train split.
eval.jsonl — first 500 rows of the source test split.
train.tiny.jsonl / eval.tiny.jsonl — 16 / 8 rows, for a fast CPU wiring dry-run.
Each row keeps five fields from the source: id, domain, sql_prompt (the… See the full description on the dataset page: https://huggingface.co/datasets/ml-infra-toloka/synthetic-text-to-sql-tangle.text-to-sql-dataset
English Text‑to‑SQL with Optional Schema
This dataset maps English user questions to SQL queries, with or without a compact database schema provided. The schema, when present, is a minimal representation of the database structure: a comma‑separated list of table names followed by their column names in parentheses. No data types are included.
Dataset Structure
Column
Type
Description
text
string
The user's request in English.
schema
string (optional)… See the full description on the dataset page: https://huggingface.co/datasets/sirunchained/text-to-sql-dataset.sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/dipanjanS/sql-create-context.sql-create-context-thai
Overview
This dataset builds from sql-create-context.
@misc{b-mc2_2023_sql-create-context,
title = {sql-create-context Dataset},
author = {b-mc2},
year = {2023},
url = {https://huggingface.co/datasets/b-mc2/sql-create-context},
note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.},
}
beaver-dw-plan-sql
BEAVER-dw Plan→SQL
A restructuring of the dw subset of BEAVER
into a plan-then-SQL format, with family-disjoint splits.
Each example asks a model to emit a structured plan first and the SQL second:
{
"question": "Which departments offered the most subjects last term?",
"domain_knowledge": ["..."],
"ir": {
"tables": ["SIS_DEPARTMENT", "SUBJECT_OFFERED_SUMMARY"],
"join_keys": [{"left": "SIS_DEPARTMENT.DEPARTMENT_CODE",
"right":… See the full description on the dataset page: https://huggingface.co/datasets/mercurylabs-ai/beaver-dw-plan-sql.text-to-sql-struct-distillation-minidev
结构化 Text-to-SQL 蒸馏 Mini-Dev 派生 SFT
本仓库发布由 BIRD Mini-Dev 500 个样本构造的派生 SFT messages 数据,共 500 条。
文件与边界
minidev_sft_messages.jsonl:500 条 messages 格式的派生 SFT 样本。
不含 Mini-Dev SQLite 数据库、gold SQL、原始 schema 文件或上游数据库内容。
评测中使用 BIRD 作者提供的 SQLite 集合型 Execution Accuracy (EX) 定义;该评测代码与复现说明在 GitHub 工程中维护。
来源、署名与许可证
本数据为 BIRD Mini-Dev 的派生文本内容。上游仓库:https://github.com/bird-bench/mini_dev。上游 README 标明 CC BY-SA 4.0,因此本仓库按 CC BY-SA 4.0 发布。使用或再发布时请保留对 BIRD 与 Mini-Dev… See the full description on the dataset page: https://huggingface.co/datasets/craboy4/text-to-sql-struct-distillation-minidev.
