Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.8k downloads10mo agoHugging Face02xu3kev /BIRD-SQL-data-train Dataset Card for "BIRD-SQL-data-train" Data from BIRD-SQL benchmark training set. text1K<n<10K16 likes2.7k downloads3y agoHugging Face03trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes901 downloads7mo agoHugging Face04koookiy /BIRD-SQL-data-train-CoTPart of BIRD sql train dataset added with chain of thought distilled from DeepSeek-R1 text1K<n<10K1 likes792 downloads2y agoHugging Face05meowterspace45 /bird-sql-train-with-reasoning bird-sql-train-with-reasoning Short description: BirdSQL training set (Text-to-SQL) enhanced with chain-of-thought / reasoning traces. License: Apache-2.0 Enhanced with reasoning traces using Nemo Data Designer and openai/gpt-oss-120b Schema (example fields) Each record contains fields like: db_id (string) question (string) evidence (nullable / string) SQL (string) schema (string) reasoning_trace (string or json-serialized object) quality_assessment (optional; string or… See the full description on the dataset page: https://huggingface.co/datasets/meowterspace45/bird-sql-train-with-reasoning.text1K<n<10K0 likes595 downloads1y agoHugging Face06hardikch05 /100000_text_to_sqltext10M<n<100M12 likes555 downloads3y agoHugging Face07lamini /bird_text_to_sql Dataset Card for "bird_text_to_sql" More Information needed text10K<n<100K7 likes470 downloads3y agoHugging Face08philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes421 downloads2y agoHugging Face09Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes412 downloads4mo agoHugging Face10lamini /bird_spider_train_text_to_sql Dataset Card for "bird_spider_train_text_to_sql" More Information needed text10K<n<100K5 likes314 downloads3y agoHugging Face11trl-lab /SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases Project page · Questions and SQL · Trained models · Python library · Citation This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.texttext-generation1K<n<10K1 likes279 downloads8d agoHugging Face12xu3kev /BIRD-SQL-data Dataset Card for "BIRD-SQL-data" More Information needed textn<1K1 likes264 downloads3y agoHugging Face13Lots-of-LoRAs /task076_splash_correcting_sql_mistake Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task076_splash_correcting_sql_mistake Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task076_splash_correcting_sql_mistake.texttext-generation1K<n<10K0 likes249 downloads2y agoHugging Face14567-labs /bird-sql-train-evaltext1K<n<10K0 likes241 downloads2y agoHugging Face15sqlrooms /earthquakes California Earthquakes 1967-2018 Location, magnitude and type of 2.5+ magnitude earthquakes in California from 1967 to 2018. Source: https://kepler.gl/ tabular10K<n<100K0 likes222 downloads1y agoHugging Face16trl-lab /SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL Project page · Schemas and databases · Trained models · Python library · Citation SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.texttext-generation100K<n<1M1 likes212 downloads8d agoHugging Face17ChrisHayduk /Llama-2-SQL-Dataset Dataset Card for "Llama-2-SQL-Dataset" This dataset is deprecated in favor of ChrisHayduk/Llama-2-SQL-and-Code-Dataset text10K<n<100K13 likes191 downloads3y agoHugging Face18PipableAI /pip-txt-to-sql-spider-bird-dataset Dataset Card for "spider-bird" More Information needed text10K<n<100K11 likes190 downloads3y agoHugging Face19lamini /spider_text_to_sql Dataset Card for "spider_text_to_sql" More Information needed text1K<n<10K9 likes174 downloads3y agoHugging Face20philikai /SQL_Spider_DDLtext1K<n<10K2 likes172 downloads3y agoHugging Face21Sudnya /bird-sql BIRD-SQL Dataset BIRD (BIg Bench for LaRge-scale Database Grounded Text-to-SQL Evaluation) is a comprehensive text-to-SQL dataset featuring realistic databases and complex queries across multiple domains. This dataset maintains the exact original BIRD format and field names. Dataset Statistics Train: ~9,400 examples Validation: ~1,500 examples Total: ~10,900 examples Databases: 80+ realistic databases Quick Start from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/bird-sql.text10K<n<100K0 likes162 downloads1y agoHugging Face22VPCSinfo /odoo-sql-query-dataset Odoo SQL Query Dataset This dataset contains natural language to SQL query pairs specifically for Odoo 17.0 Community Edition. It's designed to help train and fine-tune language models for generating accurate SQL queries for Odoo databases. Dataset Description Overview The dataset consists of 6815 carefully curated examples of natural language questions paired with their corresponding SQL queries for Odoo databases. Each example includes detailed instructions… See the full description on the dataset page: https://huggingface.co/datasets/VPCSinfo/odoo-sql-query-dataset.texttext-generation1K<n<10K5 likes137 downloads2y agoHugging Face23koookiy /BIRD-SQL-data-traintext1K<n<10K0 likes124 downloads2y agoHugging Face24Anna4242 /sql-multiturn-training-dataset-combinedtext1M<n<10M0 likes121 downloads1y agoHugging Face25chrisjcc /text-to-sql-spider-dataset Text-to-SQL Dataset A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning. 📊 Dataset Summary Total Samples: 20000 Format: Chat template (system/user/assistant messages) Task: Text-to-SQL generation Language: English License: apache-2.0 📁 Dataset Structure Data Format Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.texttext-generation10K<n<100K1 likes114 downloads1y agoHugging Face26Lots-of-LoRAs /task077_splash_explanation_to_sql Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task077_splash_explanation_to_sql Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task077_splash_explanation_to_sql.texttext-generation1K<n<10K0 likes104 downloads2y agoHugging Face27Rif-SQL /time-series-uk-retail-supermarket-price-data Dataset Card for UK Supermarket Time Series Data (Combined) This dataset is a consolidated version of the Time Series UK Supermarket Data originally shared by Declan McAlinden under the Apache License 2.0. It merges the five original CSV files (one per UK supermarket) into a single Parquet file:base_retail_gb_snappy.parquet Listed UK supermarket, Sainsbury, ASDA, Tesco, Morrisons & Aldi The data remains unchanged, but combining the files simplifies loading and analysis, especially… See the full description on the dataset page: https://huggingface.co/datasets/Rif-SQL/time-series-uk-retail-supermarket-price-data.text1M<n<10M1 likes86 downloads11mo agoHugging Face28benjamintli /BIRD-SQL-data-train-formattedtext1K<n<10K0 likes85 downloads1y agoHugging Face29xiaobing11 /ACE-SQL ACE-SQL Training Data This repository contains the curated supervised fine-tuning (SFT), reinforcement learning (RL), and empirical-pool data released with ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for Text-to-SQL. ACE-SQL trains a shared language-model policy in two roles: a schema retriever that selects the minimum required database columns, and a SQL generator that operates on the resulting pruned schema. The SFT data provides a cold start for both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.texttext-generation10K<n<100K1 likes84 downloads4mo agoHugging Face30DanielRegaladoCardoso /text-to-sql-mix-v2 🔗 Part of the SQL Agent LLMOps project This dataset is one of three purpose-built training mixes for the SQL Agent LLMOps project — an end-to-end pipeline that converts natural-language questions into SQL, executes the query on user data, and renders a storytelling-grade visualization. Dataset Model trained Role 🤗 DanielRegaladoCardoso/text-to-sql-mix-v2 Qwen 2.5 Coder 7B NL question → SQL 🤗 DanielRegaladoCardoso/chart-reasoning-mix-v1Phi-3 Mini 3.8B (question +… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/text-to-sql-mix-v2.texttext-generation100K<n<1M0 likes79 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.