Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes4.3k downloads3y agoHugging Face02gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads10mo agoHugging Face03xu3kev /BIRD-SQL-data-train Dataset Card for "BIRD-SQL-data-train" Data from BIRD-SQL benchmark training set. text1K<n<10K16 likes2.5k downloads3y agoHugging Face04birdsql /bird-critic-1.0-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. The schema file is included in the code repository https://github.com/bird-bench/BIRD-CRITIC-1/blob/main/baseline/data/sqlite_schema.jsonl BIRD-CRITIC-1.0-SQLite BIRD-Critic is the first SQL debugging… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-sqlite.textn<1K2 likes2.3k downloads7mo agoHugging Face05birdsql /bird_sql_dev_20251106 BIRD-SQL Dev 🆕 Update 2025-11-06 We would like to express our sincere gratitude to the community for their continuous support and constructive feedback on the BIRD-SQL Dev dataset. Over the past year, we have received valuable suggestions through GitHub discussions, emails, and user reports. Based on these insights, we organized a quality review program led by a team of five PhD researchers in Data Science and AI, supported by a globally distributed group of industry… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird_sql_dev_20251106.texttable-question-answering1K<n<10K10 likes1.2k downloads9mo agoHugging Face06trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes886 downloads7mo agoHugging Face07koookiy /BIRD-SQL-data-train-CoTPart of BIRD sql train dataset added with chain of thought distilled from DeepSeek-R1 text1K<n<10K1 likes731 downloads2y agoHugging Face08birdsql /livesqlbench-base-lite-sqlite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.texttable-question-answeringn<1K4 likes726 downloads4mo agoHugging Face09hardikch05 /100000_text_to_sqltext10M<n<100M12 likes625 downloads3y agoHugging Face10meowterspace45 /bird-sql-train-with-reasoning bird-sql-train-with-reasoning Short description: BirdSQL training set (Text-to-SQL) enhanced with chain-of-thought / reasoning traces. License: Apache-2.0 Enhanced with reasoning traces using Nemo Data Designer and openai/gpt-oss-120b Schema (example fields) Each record contains fields like: db_id (string) question (string) evidence (nullable / string) SQL (string) schema (string) reasoning_trace (string or json-serialized object) quality_assessment (optional; string or… See the full description on the dataset page: https://huggingface.co/datasets/meowterspace45/bird-sql-train-with-reasoning.text1K<n<10K0 likes536 downloads1y agoHugging Face11philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes465 downloads2y agoHugging Face12lamini /bird_text_to_sql Dataset Card for "bird_text_to_sql" More Information needed text10K<n<100K7 likes462 downloads3y agoHugging Face13while-ai /text-to-sql-shop Text-to-SQL on a seeded store schema, with checkpoints Recipe: recipes/04-train/text-to-sql · Collection: Analyst A question about an online store's database in, one PostgreSQL query out, graded by a program: run the query, compare the result set to the gold query's result. The schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and the benchmark runner are the recipes/04-train/text-to-sql recipe in the open-source whileai SDK. Splits |… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.tabular10K<n<100K0 likes454 downloads14d agoHugging Face14Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes345 downloads4mo agoHugging Face15birdsql /six-gym-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. This dataset is the train split of BIRD-Critic-SQLite, comprising 5,000 data instances for model training and development. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. 📋 Dataset Structure Below is a description of the dataset fields and additional… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-sqlite.text1K<n<10K0 likes326 downloads7mo agoHugging Face16ibm-research /SQL-API-Bench Dataset Card for Dataset Name This dataset contains QA that requires DB and API access at the same time. It is composed of two new benchmarks consisting of questions whose answers require a combination of database and API calls, both of which are augmentations of the popular Spider dataset and benchmark. Benchmark I replaces a fraction of the real Spider database tables with equivalents that are executed via APIs. This allows us to directly test the mechanism by which database and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SQL-API-Bench.textquestion-answering1K<n<10K5 likes323 downloads1y agoHugging Face17Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes321 downloads3y agoHugging Face18lamini /bird_spider_train_text_to_sql Dataset Card for "bird_spider_train_text_to_sql" More Information needed text10K<n<100K5 likes309 downloads3y agoHugging Face19rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes285 downloads1mo agoHugging Face20xu3kev /BIRD-SQL-data Dataset Card for "BIRD-SQL-data" More Information needed textn<1K1 likes263 downloads3y agoHugging Face21Lots-of-LoRAs /task076_splash_correcting_sql_mistake Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task076_splash_correcting_sql_mistake Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task076_splash_correcting_sql_mistake.texttext-generation1K<n<10K0 likes263 downloads2y agoHugging Face22mlfoundations-dev /b2_code_fasttext_pos_ioi_neg_sql_eval_636d mlfoundations-dev/b2_code_fasttext_pos_ioi_neg_sql_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 16.7 57.8 76.2 27.8 39.4 43.4 41.9 14.9 17.7 AIME24 Average Accuracy: 16.67% ± 1.33% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30 2 13.33% 4 30 3 16.67% 5 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_ioi_neg_sql_eval_636d.tabular1K<n<10K0 likes252 downloads1y agoHugging Face23trl-lab /SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases Project page · Questions and SQL · Trained models · Python library · Citation This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.texttext-generation1K<n<10K0 likes229 downloads4d agoHugging Face24sqlrooms /earthquakes California Earthquakes 1967-2018 Location, magnitude and type of 2.5+ magnitude earthquakes in California from 1967 to 2018. Source: https://kepler.gl/ tabular10K<n<100K0 likes228 downloads1y agoHugging Face25567-labs /bird-sql-train-evaltext1K<n<10K0 likes224 downloads2y agoHugging Face26lamini /spider_text_to_sql Dataset Card for "spider_text_to_sql" More Information needed text1K<n<10K9 likes218 downloads3y agoHugging Face27Anna4242 /sql-multiturn-training-dataset-combinedtext1M<n<10M0 likes199 downloads1y agoHugging Face28ShHugging /BIRD-SQL-TRAINtext1K<n<10K0 likes197 downloads2y agoHugging Face29ChrisHayduk /Llama-2-SQL-Dataset Dataset Card for "Llama-2-SQL-Dataset" This dataset is deprecated in favor of ChrisHayduk/Llama-2-SQL-and-Code-Dataset text10K<n<100K13 likes194 downloads3y agoHugging Face30PipableAI /pip-txt-to-sql-spider-bird-dataset Dataset Card for "spider-bird" More Information needed text10K<n<100K11 likes183 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.