datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.BIRD-SQL-data-train
Dataset Card for "BIRD-SQL-data-train"
Data from BIRD-SQL benchmark training set.
SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.BIRD-SQL-data-train-CoTPart of BIRD sql train dataset added with chain of thought distilled from DeepSeek-R1
bird-sql-train-with-reasoning
bird-sql-train-with-reasoning
Short description: BirdSQL training set (Text-to-SQL) enhanced with chain-of-thought / reasoning traces.
License: Apache-2.0
Enhanced with reasoning traces using Nemo Data Designer and openai/gpt-oss-120b
Schema (example fields)
Each record contains fields like:
db_id (string)
question (string)
evidence (nullable / string)
SQL (string)
schema (string)
reasoning_trace (string or json-serialized object)
quality_assessment (optional; string or… See the full description on the dataset page: https://huggingface.co/datasets/meowterspace45/bird-sql-train-with-reasoning.100000_text_to_sqlbird_text_to_sql
Dataset Card for "bird_text_to_sql"
More Information needed
gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.bird_spider_train_text_to_sql
Dataset Card for "bird_spider_train_text_to_sql"
More Information needed
SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases
Project page ·
Questions and SQL ·
Trained models ·
Python library ·
Citation
This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.BIRD-SQL-data
Dataset Card for "BIRD-SQL-data"
More Information needed
task076_splash_correcting_sql_mistake
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task076_splash_correcting_sql_mistake
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task076_splash_correcting_sql_mistake.bird-sql-train-evalearthquakes
California Earthquakes 1967-2018
Location, magnitude and type of 2.5+ magnitude earthquakes in California from 1967 to 2018.
Source: https://kepler.gl/
SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL
Project page ·
Schemas and databases ·
Trained models ·
Python library ·
Citation
SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.Llama-2-SQL-Dataset
Dataset Card for "Llama-2-SQL-Dataset"
This dataset is deprecated in favor of ChrisHayduk/Llama-2-SQL-and-Code-Dataset
pip-txt-to-sql-spider-bird-dataset
Dataset Card for "spider-bird"
More Information needed
spider_text_to_sql
Dataset Card for "spider_text_to_sql"
More Information needed
SQL_Spider_DDLbird-sql
BIRD-SQL Dataset
BIRD (BIg Bench for LaRge-scale Database Grounded Text-to-SQL Evaluation) is a comprehensive text-to-SQL dataset featuring realistic databases and complex queries across multiple domains.
This dataset maintains the exact original BIRD format and field names.
Dataset Statistics
Train: ~9,400 examples
Validation: ~1,500 examples
Total: ~10,900 examples
Databases: 80+ realistic databases
Quick Start
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/bird-sql.odoo-sql-query-dataset
Odoo SQL Query Dataset
This dataset contains natural language to SQL query pairs specifically for Odoo 17.0 Community Edition. It's designed to help train and fine-tune language models for generating accurate SQL queries for Odoo databases.
Dataset Description
Overview
The dataset consists of 6815 carefully curated examples of natural language questions paired with their corresponding SQL queries for Odoo databases. Each example includes detailed instructions… See the full description on the dataset page: https://huggingface.co/datasets/VPCSinfo/odoo-sql-query-dataset.BIRD-SQL-data-trainsql-multiturn-training-dataset-combinedtext-to-sql-spider-dataset
Text-to-SQL Dataset
A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning.
📊 Dataset Summary
Total Samples: 20000
Format: Chat template (system/user/assistant messages)
Task: Text-to-SQL generation
Language: English
License: apache-2.0
📁 Dataset Structure
Data Format
Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.task077_splash_explanation_to_sql
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task077_splash_explanation_to_sql
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task077_splash_explanation_to_sql.time-series-uk-retail-supermarket-price-data
Dataset Card for UK Supermarket Time Series Data (Combined)
This dataset is a consolidated version of the Time Series UK Supermarket Data originally shared by Declan McAlinden under the Apache License 2.0.
It merges the five original CSV files (one per UK supermarket) into a single Parquet file:base_retail_gb_snappy.parquet
Listed UK supermarket, Sainsbury, ASDA, Tesco, Morrisons & Aldi
The data remains unchanged, but combining the files simplifies loading and analysis, especially… See the full description on the dataset page: https://huggingface.co/datasets/Rif-SQL/time-series-uk-retail-supermarket-price-data.BIRD-SQL-data-train-formattedACE-SQL
ACE-SQL Training Data
This repository contains the curated supervised fine-tuning (SFT),
reinforcement learning (RL), and empirical-pool data released with
ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for
Text-to-SQL.
ACE-SQL trains a shared language-model policy in two roles: a schema retriever
that selects the minimum required database columns, and a SQL generator that
operates on the resulting pruned schema. The SFT data provides a cold start for
both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.text-to-sql-mix-v2
🔗 Part of the SQL Agent LLMOps project
This dataset is one of three purpose-built training mixes for the
SQL Agent LLMOps project — an end-to-end pipeline that
converts natural-language questions into SQL, executes the query on user
data, and renders a storytelling-grade visualization.
Dataset
Model trained
Role
🤗 DanielRegaladoCardoso/text-to-sql-mix-v2
Qwen 2.5 Coder 7B
NL question → SQL
🤗 DanielRegaladoCardoso/chart-reasoning-mix-v1Phi-3 Mini 3.8B
(question +… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/text-to-sql-mix-v2.
