Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes4.3k downloads3y agoHugging Face02gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads10mo agoHugging Face03trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes902 downloads7mo agoHugging Face04philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes466 downloads2y agoHugging Face05Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes344 downloads4mo agoHugging Face06Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes321 downloads3y agoHugging Face07Lots-of-LoRAs /task076_splash_correcting_sql_mistake Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task076_splash_correcting_sql_mistake Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task076_splash_correcting_sql_mistake.texttext-generation1K<n<10K0 likes263 downloads2y agoHugging Face08trl-lab /SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases Project page · Questions and SQL · Trained models · Python library · Citation This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.texttext-generation1K<n<10K0 likes231 downloads4d agoHugging Face09hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes179 downloads13d agoHugging Face10trl-lab /SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL Project page · Schemas and databases · Trained models · Python library · Citation SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.texttext-generation100K<n<1M0 likes166 downloads4d agoHugging Face11emgena /omnimcp_sql_postgres_pro_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_postgres_pro_teaser.text-generationn<1K0 likes158 downloads19d agoHugging Face12jumplander /Persian-Business-Text-to-SQL-Gold-1K Persian Business Text-to-SQL Gold-1K 1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking. مجموعه‌ای ۱۰۰۰ نمونه‌ای برای تبدیل درخواست‌های فارسی کسب‌وکار به SQL، همراه با دیتابیس‌های SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy. Motivation BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.texttext-generation1K<n<10K2 likes157 downloads1mo agoHugging Face13chrisjcc /text-to-sql-spider-dataset Text-to-SQL Dataset A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning. 📊 Dataset Summary Total Samples: 20000 Format: Chat template (system/user/assistant messages) Task: Text-to-SQL generation Language: English License: apache-2.0 📁 Dataset Structure Data Format Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.texttext-generation10K<n<100K1 likes151 downloads1y agoHugging Face14emgena /omnimcp_sql_snowflake_warehouse_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_snowflake_warehouse_teaser.text-generationn<1K0 likes141 downloads19d agoHugging Face15emgena /omnimcp_sql_dbt_transformations_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_dbt_transformations_teaser.text-generationn<1K0 likes138 downloads19d agoHugging Face16Lots-of-LoRAs /task077_splash_explanation_to_sql Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task077_splash_explanation_to_sql Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task077_splash_explanation_to_sql.texttext-generation1K<n<10K0 likes124 downloads2y agoHugging Face17VPCSinfo /odoo-sql-query-dataset Odoo SQL Query Dataset This dataset contains natural language to SQL query pairs specifically for Odoo 17.0 Community Edition. It's designed to help train and fine-tune language models for generating accurate SQL queries for Odoo databases. Dataset Description Overview The dataset consists of 6815 carefully curated examples of natural language questions paired with their corresponding SQL queries for Odoo databases. Each example includes detailed instructions… See the full description on the dataset page: https://huggingface.co/datasets/VPCSinfo/odoo-sql-query-dataset.texttext-generation1K<n<10K5 likes121 downloads2y agoHugging Face18KBlueLeaf /Danbooru2021-SQLite Danbooru 2021 SQLite Dataset Summary This is the metadata of danbooru 2021 dataset in SQLite format. https://gwern.net/danbooru2021 Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/Danbooru2021-SQLite.text-generation1M<n<10M10 likes106 downloads4y agoHugging Face19adelelsayed1991 /fhirsql-reasoning-sql FHIR-to-SQL with Database-Resolved Clinical Terminology Natural-language hospital questions paired with a structured JSON query plan and compiled DuckDB SQL, over a FHIR-derived clinical schema. Built entirely from synthetic Synthea patients — no real patient data. Companion to the paper Plan-Then-Compile: Turning a General-Purpose Coder Model into a FHIR Data Analyst. Code & paper: https://github.com/adelelsayed/fhirsql-reasoning-sql Adapters:… See the full description on the dataset page: https://huggingface.co/datasets/adelelsayed1991/fhirsql-reasoning-sql.text-generation10K<n<100K0 likes104 downloads1mo agoHugging Face20cstr /de-wiktionary-sqlite-normalized German Wiktionary - Normalized SQLite Database A fully normalized, production-ready SQLite database of German Wiktionary with complete linguistic information and optimized query performance. 🎯 Key Features ✅ Zero data loss: All information from original Wiktionary preserved ⚡ Lightning-fast queries: Comprehensive indexing (< 5ms typical queries) 🔍 Full grammatical analysis: Complete inflection paradigms, word forms, 185 unique grammatical tags 🔗 Semantic relations:… See the full description on the dataset page: https://huggingface.co/datasets/cstr/de-wiktionary-sqlite-normalized.text-retrieval100K<n<1M1 likes97 downloads11mo agoHugging Face21sigdelakshey /SQLRobustBench SQLRobustBench SQLRobustBench is a synthetic benchmark for SQL robustness under explicit schema, parsing, and logical validation constraints. Dataset Summary This release focuses on two benchmark families: SQLCorrupt: invalid SQL detection and repair SQLNormalize: deterministic SQL canonicalization and normalization SQLRobustBench evaluates model behavior on SQL tasks that require more than straightforward generation: generating valid clean SQL under schema constraints… See the full description on the dataset page: https://huggingface.co/datasets/sigdelakshey/SQLRobustBench.text-generation1K<n<10K1 likes95 downloads7mo agoHugging Face22xiaobing11 /ACE-SQL ACE-SQL Training Data This repository contains the curated supervised fine-tuning (SFT), reinforcement learning (RL), and empirical-pool data released with ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for Text-to-SQL. ACE-SQL trains a shared language-model policy in two roles: a schema retriever that selects the minimum required database columns, and a SQL generator that operates on the resulting pruned schema. The SFT data provides a cold start for both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.texttext-generation10K<n<100K1 likes92 downloads4mo agoHugging Face23hiimivantang /bird-sqlite-sft-train BIRD SQLite Text-to-SQL SFT Dataset Supervised fine-tuning data for a SQLite-dialect text-to-SQL specialist model, built from the BIRD benchmark train split. Contents 7,483 train + 408 val examples spanning 69 distinct database schemas (movie_platform, chicago_crime, hockey, mondial_geo, works_cycles, and 64 others), split by a stratified per-database 95/5 hold-out (sft_sqlite_ train.jsonl / sft_sqlite_val.jsonl) with zero exact overlap between them. Format:… See the full description on the dataset page: https://huggingface.co/datasets/hiimivantang/bird-sqlite-sft-train.texttext-generation1K<n<10K0 likes91 downloads16d agoHugging Face24Lots-of-LoRAs /task107_splash_question_to_sql Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task107_splash_question_to_sql Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task107_splash_question_to_sql.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face25birdsql /effi-sql-training Diff-SQL Training Dataset We release the training datasets used by Diff-SQL for SQL efficiency optimization. This dataset includes: Patch Generator Training Dataset: SFT data for generating SQL optimization patches. Constraint Aligner Training Dataset: SFT warmup data for constraint-aware SQL optimization refinement. Files patch-generator-training-dataset/ train.parquet dev.parquet constraint-aligner-training-dataset/ train.parquet dev.parquet… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/effi-sql-training.texttext-generation1K<n<10K0 likes88 downloads4mo agoHugging Face26philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes86 downloads3y agoHugging Face27salmane11 /SQLShield SQLShield Dataset Summary SQLShield is a dataset designed for training and evaluating models on detecting vulnerable versus benign SQL usage in natural language-driven database interfaces. It includes a rich collection of natural language questions, their corresponding SQL queries, relevant table contexts, and a binary vulnerability label indicating whether the SQL query is potentially malicious (1) or safe (0). This dataset enables research to improve safety in… See the full description on the dataset page: https://huggingface.co/datasets/salmane11/SQLShield.texttext-generation10K<n<100K1 likes75 downloads1y agoHugging Face28hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes73 downloads19d agoHugging Face29emdemor /sql-create-context-pt Overview Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context, que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas utilizando a instrução CREATE TABLE como contexto. O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.texttext-generation10K<n<100K2 likes71 downloads2y agoHugging Face30birdsql /Effi-SQL Effi-SQL Update 2026-06-12 We release Effi-SQL, a dataset suite for SQL efficiency optimization. This collection includes: Effi-SQL Benchmark: a benchmark for evaluating SQL efficiency optimization methods. Diff-SQL Training Dataset: training data used by Diff-SQL, including data for the Patch Generator and Constraint Aligner. Dataset Fields Effi-SQL Benchmark id: A unique identifier for each benchmark instance. db: The database… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/Effi-SQL.tabulartext-generationn<1K0 likes69 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.