Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.8k downloads10mo agoHugging Face02trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes901 downloads7mo agoHugging Face03hardikch05 /100000_text_to_sqltext10M<n<100M12 likes555 downloads3y agoHugging Face04while-ai /text-to-sql-shop Text-to-SQL on a seeded store schema, with checkpoints Recipe: recipes/04-train/text-to-sql · Collection: Analyst A question about an online store's database in, one PostgreSQL query out, graded by a program: run the query, compare the result set to the gold query's result. The schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and the benchmark runner are the recipes/04-train/text-to-sql recipe in the open-source whileai SDK. Splits |… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.tabular10K<n<100K0 likes496 downloads18d agoHugging Face05lamini /bird_text_to_sql Dataset Card for "bird_text_to_sql" More Information needed text10K<n<100K7 likes470 downloads3y agoHugging Face06philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes421 downloads2y agoHugging Face07Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes412 downloads4mo agoHugging Face08Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes355 downloads3y agoHugging Face09lamini /bird_spider_train_text_to_sql Dataset Card for "bird_spider_train_text_to_sql" More Information needed text10K<n<100K5 likes314 downloads3y agoHugging Face10trl-lab /SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases Project page · Questions and SQL · Trained models · Python library · Citation This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.texttext-generation1K<n<10K1 likes279 downloads8d agoHugging Face11trl-lab /SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL Project page · Schemas and databases · Trained models · Python library · Citation SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.texttext-generation100K<n<1M1 likes212 downloads8d agoHugging Face12lamini /spider_text_to_sql Dataset Card for "spider_text_to_sql" More Information needed text1K<n<10K9 likes174 downloads3y agoHugging Face13hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes137 downloads18d agoHugging Face14chrisjcc /text-to-sql-spider-dataset Text-to-SQL Dataset A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning. 📊 Dataset Summary Total Samples: 20000 Format: Chat template (system/user/assistant messages) Task: Text-to-SQL generation Language: English License: apache-2.0 📁 Dataset Structure Data Format Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.texttext-generation10K<n<100K1 likes114 downloads1y agoHugging Face15jumplander /Persian-Business-Text-to-SQL-Gold-1K Persian Business Text-to-SQL Gold-1K 1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking. مجموعه‌ای ۱۰۰۰ نمونه‌ای برای تبدیل درخواست‌های فارسی کسب‌وکار به SQL، همراه با دیتابیس‌های SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy. Motivation BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.texttext-generation1K<n<10K2 likes95 downloads1mo agoHugging Face16Clinton /texttosqlv2_25000_v2text10K<n<100K8 likes81 downloads3y agoHugging Face17DanielRegaladoCardoso /text-to-sql-mix-v2 🔗 Part of the SQL Agent LLMOps project This dataset is one of three purpose-built training mixes for the SQL Agent LLMOps project — an end-to-end pipeline that converts natural-language questions into SQL, executes the query on user data, and renders a storytelling-grade visualization. Dataset Model trained Role 🤗 DanielRegaladoCardoso/text-to-sql-mix-v2 Qwen 2.5 Coder 7B NL question → SQL 🤗 DanielRegaladoCardoso/chart-reasoning-mix-v1Phi-3 Mini 3.8B (question +… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/text-to-sql-mix-v2.texttext-generation100K<n<1M0 likes79 downloads6mo agoHugging Face18hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes75 downloads24d agoHugging Face19MrezaPRZ /bird_text_to_sqltext1K<n<10K0 likes68 downloads2y agoHugging Face20hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes67 downloads23d agoHugging Face21Yihim /synthetic_text_to_sql-qwen2.5-instruct-curatedtext100K<n<1M1 likes57 downloads2y agoHugging Face22NickyNicky /synthetic_text_to_sql_format_chatML_gemma dataset base. gretelai/synthetic_text_to_sql dataset = load_dataset("NickyNicky/synthetic_text_to_sql_format_chatML_gemma") <bos><start_of_turn>system You are a helpful AI assistant. you are a sql expert who responds in json format.<end_of_turn> <start_of_turn>user ## prompt: What is the total gold production by 'Site B' in the 'production' table? ## sql context: CREATE TABLE production (id INT, site VARCHAR(50), year INT, gold_production INT, silver_production… See the full description on the dataset page: https://huggingface.co/datasets/NickyNicky/synthetic_text_to_sql_format_chatML_gemma.text100K<n<1M1 likes54 downloads3y agoHugging Face23VictorDCh /spider-clean-text-to-sqltext1K<n<10K1 likes51 downloads2y agoHugging Face24nnikolovskii /text-to-sql-validationtext10K<n<100K0 likes50 downloads2y agoHugging Face25FatimahEmadEldin /MycoBase-Large-Scale-Text-to-SQL MycoBase: A Biologically Literate Text-to-SQL Dataset MycoBase is a synthetic but biologically accurate dataset designed for stress-testing Text-to-SQL systems. It represents a research information system for the study of fungi, covering everything from taxonomy and genomics to morphology and cultivation. Dataset Highlights Schema Complexity: 2,016 tables with over 9,000 foreign key relationships. Data Volume: 320,270 rows of realistic mycology data. Realistic Names:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/MycoBase-Large-Scale-Text-to-SQL.tabular100K<n<1M0 likes47 downloads5mo agoHugging Face26Safreliy /synthetic_text_to_sqltexttext-generation10K<n<100K0 likes40 downloads1y agoHugging Face27VictorDCh /spider-clean-text-to-sql-2text1K<n<10K0 likes39 downloads2y agoHugging Face28meowterspace45 /synthetic_text_to_sql_reasoning Synthetic Text-to-SQL with Reasoning Traces This dataset is an enhanced version of gretelai/synthetic_text_to_sql with synthetic reasoning traces added using Nemo Data Designer and openai/gpt-oss-120b for generation. Dataset Description gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. The original dataset includes: 105,851 records partitioned into… See the full description on the dataset page: https://huggingface.co/datasets/meowterspace45/synthetic_text_to_sql_reasoning.texttext-generation10K<n<100K1 likes38 downloads1y agoHugging Face29AmanPriyanshu /reasoning-sft-synthetic_text_to_sql-128K synthetic_text_to_sql (converted) Converted version of gretelai/synthetic_text_to_sql, reformatted to 100,000 rows for reasoning SFT training. Format Each row has three columns: input — list of dicts [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}] (system prompt contains the database schema, user prompt contains the natural language question) response — response string with <think> reasoning block (SQL explanation) followed by the SQL query… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-synthetic_text_to_sql-128K.textquestion-answering100K<n<1M0 likes36 downloads7mo agoHugging Face30TafcoMetawireless /synthetic_text_to_sql_en_es Dataset basado en la versión de GretelAI - SyntheticSQL synthetic_text_to_sql_en_es Se trata de una expansión mediante la traducción al español de la columna 'sql_prompt'. Se ha añadido una columna extra 'sql_prompt_es' que contiene el prompt original de inglés traducido al español. Para obtener estas traducciones, se utilizó few-shot prompting + CoT mediante el modelo Qwen/Qwen2.5-32B-Instruct-AWQ Actualización 6/27/25 En la versión pasada se encontraron… See the full description on the dataset page: https://huggingface.co/datasets/TafcoMetawireless/synthetic_text_to_sql_en_es.textquestion-answering100K<n<1M3 likes35 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.