datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.100000_text_to_sqlgretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
bird_text_to_sql
Dataset Card for "bird_text_to_sql"
More Information needed
text-to-sql-shop
Text-to-SQL on a seeded store schema, with checkpoints
Recipe: recipes/04-train/text-to-sql · Collection: Analyst
A question about an online store's database in, one PostgreSQL query out, graded by
a program: run the query, compare the result set to the gold query's result. The
schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and
the benchmark runner are the
recipes/04-train/text-to-sql
recipe in the open-source whileai SDK.
Splits
|… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.Text-to-sql-v1bird_spider_train_text_to_sql
Dataset Card for "bird_spider_train_text_to_sql"
More Information needed
text-to-sql-with-evidence-training-pool
Text to SQL training pool with evidence
A training pool for text-to-SQL models, assembled from the public corpora that publish both a
question and a SQL query, with the databases those queries run against. Every query here was
executed against its own database when the pool was built, and the rows whose query did not run
were left out, so a query in this pool is a query that works.
One of the sources publishes an evidence sentence with each question, the outside knowledge the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/text-to-sql-with-evidence-training-pool.SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases
Project page ·
Questions and SQL ·
Trained models ·
Python library ·
Citation
This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.spider_text_to_sql
Dataset Card for "spider_text_to_sql"
More Information needed
enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,087 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a
self-correction loop — and the benchmark is designed so that number cannot be
inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL
Project page ·
Schemas and databases ·
Trained models ·
Python library ·
Citation
SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.text-to-sql-spider-dataset
Text-to-SQL Dataset
A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning.
📊 Dataset Summary
Total Samples: 20000
Format: Chat template (system/user/assistant messages)
Task: Text-to-SQL generation
Language: English
License: apache-2.0
📁 Dataset Structure
Data Format
Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.text-to-sql-training-pool
Text to SQL training pool
A training pool for text-to-SQL models, assembled from the public corpora that publish both a
question and a SQL query, with the databases those queries run against. Every query here was
executed against its own database when the pool was built, and the rows whose query did not run
were left out, so a query in this pool is a query that works.
The pool ships in two layers over the same data. Use whichever suits the job.
rows.jsonl is the normalised… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/text-to-sql-training-pool.Rethinking_Enterprise_Text_to_SQL
Rethinking Enterprise Text to SQL Dataset
📖 Overview
This repository serves as the foundational data warehouse for an Enterprise-grade Vertical Domain Question-Answering Dataset.
Unlike traditional Text-to-SQL benchmarks, this dataset simulates real-world enterprise scenarios where users possess only a natural language question and a large, unstructured knowledge base (long_doc). The core challenge lies in knowledge mining and reasoning: models must extract… See the full description on the dataset page: https://huggingface.co/datasets/XuWave/Rethinking_Enterprise_Text_to_SQL.spider-clean-text-to-sqltext-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.bird_text_to_sqltext-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.synthetic_text_to_sql_format_chatML_gemma
dataset base.
gretelai/synthetic_text_to_sql
dataset = load_dataset("NickyNicky/synthetic_text_to_sql_format_chatML_gemma")
<bos><start_of_turn>system
You are a helpful AI assistant.
you are a sql expert who responds in json format.<end_of_turn>
<start_of_turn>user
## prompt:
What is the total gold production by 'Site B' in the 'production' table?
## sql context:
CREATE TABLE production (id INT, site VARCHAR(50), year INT, gold_production INT, silver_production… See the full description on the dataset page: https://huggingface.co/datasets/NickyNicky/synthetic_text_to_sql_format_chatML_gemma.synthetic_text_to_sql-qwen2.5-instruct-curatedtexttosqlv2_25000_v2repro-beyond-text-to-sql-can-llms-really-debug-enterprise-etl-sql-traces
Agent traces
Agent sessions published from a Trackio Logbook.
text-to-sql-validationMycoBase-Large-Scale-Text-to-SQL
MycoBase: A Biologically Literate Text-to-SQL Dataset
MycoBase is a synthetic but biologically accurate dataset designed for stress-testing Text-to-SQL systems. It represents a research information system for the study of fungi, covering everything from taxonomy and genomics to morphology and cultivation.
Dataset Highlights
Schema Complexity: 2,016 tables with over 9,000 foreign key relationships.
Data Volume: 320,270 rows of realistic mycology data.
Realistic Names:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/MycoBase-Large-Scale-Text-to-SQL.synthetic_chat_text_to_sqlText-to-sql-v1
