Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes4.1k downloads3y agoHugging Face02birdsql /bird-critic-1.0-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. The schema file is included in the code repository https://github.com/bird-bench/BIRD-CRITIC-1/blob/main/baseline/data/sqlite_schema.jsonl BIRD-CRITIC-1.0-SQLite BIRD-Critic is the first SQL debugging… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-sqlite.textn<1K2 likes2k downloads7mo agoHugging Face03birdsql /bird_sql_dev_20251106 BIRD-SQL Dev 🆕 Update 2025-11-06 We would like to express our sincere gratitude to the community for their continuous support and constructive feedback on the BIRD-SQL Dev dataset. Over the past year, we have received valuable suggestions through GitHub discussions, emails, and user reports. Based on these insights, we organized a quality review program led by a team of five PhD researchers in Data Science and AI, supported by a globally distributed group of industry… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird_sql_dev_20251106.texttable-question-answering1K<n<10K10 likes1.2k downloads9mo agoHugging Face04birdsql /livesqlbench-base-lite-sqlite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.texttable-question-answeringn<1K4 likes665 downloads5mo agoHugging Face05birdsql /six-gym-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. This dataset is the train split of BIRD-Critic-SQLite, comprising 5,000 data instances for model training and development. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. 📋 Dataset Structure Below is a description of the dataset fields and additional… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-sqlite.text1K<n<10K0 likes616 downloads7mo agoHugging Face06while-ai /text-to-sql-shop Text-to-SQL on a seeded store schema, with checkpoints Recipe: recipes/04-train/text-to-sql · Collection: Analyst A question about an online store's database in, one PostgreSQL query out, graded by a program: run the query, compare the result set to the gold query's result. The schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and the benchmark runner are the recipes/04-train/text-to-sql recipe in the open-source whileai SDK. Splits |… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.tabular10K<n<100K0 likes496 downloads18d agoHugging Face07rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes456 downloads1mo agoHugging Face08Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes355 downloads3y agoHugging Face09ibm-research /SQL-API-Bench Dataset Card for Dataset Name This dataset contains QA that requires DB and API access at the same time. It is composed of two new benchmarks consisting of questions whose answers require a combination of database and API calls, both of which are augmentations of the popular Spider dataset and benchmark. Benchmark I replaces a fraction of the real Spider database tables with equivalents that are executed via APIs. This allows us to directly test the mechanism by which database and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SQL-API-Bench.textquestion-answering1K<n<10K5 likes342 downloads1y agoHugging Face10ShHugging /BIRD-SQL-TRAINtext1K<n<10K0 likes209 downloads2y agoHugging Face11heegyu /bird-sql-mini-devtextn<1K1 likes137 downloads2y agoHugging Face12hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes137 downloads17d agoHugging Face13hiimivantang /bird-sqlite-sft-train BIRD SQLite Text-to-SQL SFT Dataset Supervised fine-tuning data for a SQLite-dialect text-to-SQL specialist model, built from the BIRD benchmark train split. Contents 7,483 train + 408 val examples spanning 69 distinct database schemas (movie_platform, chicago_crime, hockey, mondial_geo, works_cycles, and 64 others), split by a stratified per-database 95/5 hold-out (sft_sqlite_ train.jsonl / sft_sqlite_val.jsonl) with zero exact overlap between them. Format:… See the full description on the dataset page: https://huggingface.co/datasets/hiimivantang/bird-sqlite-sft-train.texttext-generation1K<n<10K0 likes134 downloads20d agoHugging Face14jumplander /Persian-Business-Text-to-SQL-Gold-1K Persian Business Text-to-SQL Gold-1K 1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking. مجموعه‌ای ۱۰۰۰ نمونه‌ای برای تبدیل درخواست‌های فارسی کسب‌وکار به SQL، همراه با دیتابیس‌های SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy. Motivation BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.texttext-generation1K<n<10K2 likes95 downloads1mo agoHugging Face15philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes80 downloads3y agoHugging Face16jtjt520j /CSpider_and_DUSQL_sql_create_context此数据集原始数据集是合并CSpider和DUSQL。根据b-mc2/sql-create-context所创建。 text10K<n<100K5 likes79 downloads3y agoHugging Face17hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes75 downloads23d agoHugging Face18k19862217 /sql-optimizer Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/k19862217/sql-optimizer.textn<1K2 likes72 downloads3y agoHugging Face19griffith-bigdata /av_sql_preprocessed_data Dataset Card for Preprocessed Text-to-SQL Benchmarks This repository contains preprocessed data for several text-to-SQL benchmarks, as presented in the paper AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views. The official code for the AV-SQL framework can be found on GitHub: pminhtam/AV-SQL. Dataset Summary This repository contains preprocessed data for several text-to-SQL benchmarks: BIRD KaggleDBQA Spider sciencebenchmark BEAVER Spider2-Lite… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/av_sql_preprocessed_data.texttable-question-answeringn<1K1 likes71 downloads3mo agoHugging Face20birdsql /Effi-SQL Effi-SQL Update 2026-06-12 We release Effi-SQL, a dataset suite for SQL efficiency optimization. This collection includes: Effi-SQL Benchmark: a benchmark for evaluating SQL efficiency optimization methods. Diff-SQL Training Dataset: training data used by Diff-SQL, including data for the Patch Generator and Constraint Aligner. Dataset Fields Effi-SQL Benchmark id: A unique identifier for each benchmark instance. db: The database… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/Effi-SQL.tabulartext-generationn<1K0 likes70 downloads4mo agoHugging Face21agentlans /sql-text-collection SQL Text Collection This is a collection of publicly available text-to-SQL datasets. Dataset Structure Each row contains the columns: context: The schema for the database (e.g., CREATE TABLE statements). query: A natural language query or action to perform, expressed in English. source: The original dataset from which the row was sourced. dialect: One or more SQL dialects identified based on dialect-specific keywords found in the context and query. If there are multiple… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/sql-text-collection.text100K<n<1M1 likes68 downloads1y agoHugging Face22hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes67 downloads22d agoHugging Face23knowrohit07 /know_sqlplease use the val ign file for training, its much cleaner. thanks :) text10K<n<100K115 likes63 downloads3y agoHugging Face24hujudev /spider-text-2-sqltext1K<n<10K0 likes63 downloads2y agoHugging Face25debugger123 /SQLFlow Text2SQL-Flow Dataset Repository This repository contains the SQLFlow dataset. The SQLFlow dataset is a large-scale, high-quality collection of semantically valid and structurally diverse Text-to-SQL examples, generated using a comprehensive SQL-aware data augmentation framework. For more details, please visit the GitHub repository:🔗 https://github.com/TechNomad-ds/Text2SQL-Flow texttext-generation10K<n<100K1 likes60 downloads8mo agoHugging Face26burtenshaw /openenv-sql-investigation OpenEnv SQL Investigation Ten original procedural SQL investigation families teach joins, duplicate-safe aggregation, missing-data handling, weighted rates, temporal conditions, cohorts, anti-joins, ranking, and streak analysis. All data are synthetic and generated locally; there are no personal data or downloaded task assets. The container generates a fresh in-memory SQLite database on every unseeded reset. tasks.jsonl contains one reproducible seed-42 instance per family, with… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/openenv-sql-investigation.textn<1K0 likes60 downloads4d agoHugging Face271digitaldesign /mirror-sql MIRROR-SQL Provenance-Controlled Database Environments for Text-to-SQL Agents. 13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs. MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus. Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources. MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.texttable-question-answeringn<1K0 likes59 downloads2mo agoHugging Face28Kaveny /sql-injection SQL注入推理能力微调数据集 概述 本数据集旨在帮助研究人员和工程师通过特定案例来微调模型在SQL注入(SQLi)检测与预防方面的能力。SQL注入是一种代码注入技术,攻击者通过将恶意的SQL查询或语句插入应用程序的输入字段中,以操纵数据库执行非授权的操作。 数据集用途 研究用途:为安全领域的研究人员提供实际案例,以便于探索和开发新的防御策略。 模型训练:为机器学习模型提供训练素材,以提高其识别和防范SQL注入攻击的能力。 教育目的:作为教育资源,帮助学生和新手了解SQL注入的风险及其防护措施。 获取更多数据 如需获取更多相关数据或希望参与贡献,请访问我们的GitHub仓库: AAuZZ/SQLiDataset 许可证 本项目使用Apache 2.0许可证。有关详细信息,请参阅LICENSE文件。 Dataset for Fine-tuning Reasoning Ability in SQL Injection Overview… See the full description on the dataset page: https://huggingface.co/datasets/Kaveny/sql-injection.text10K<n<100K6 likes58 downloads2y agoHugging Face29samrat-kar /ap-sql-peft ap-sql-peft A chat-format text-to-SQL dataset for Accounts Payable analytics on Oracle. The schema is modelled on Oracle Fusion AP, Payments and Supplier tables, and the data behind it is synthetic. The dataset was used to train the LoRA adapter samrat-kar/ap-sql-v1. Every gold SQL query was run against the demo database when the dataset was built; meta.result_rows records how many rows it returned. Format One JSON object per line: {"messages": [{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/samrat-kar/ap-sql-peft.texttext-generation1K<n<10K0 likes43 downloads14d agoHugging Face30bernabepuente /database-sql-instruction-dataset Database & SQL Instruction Dataset High-quality instruction-response pairs covering PostgreSQL, advanced queries, indexing strategies, and database optimization. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Database topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/database-sql-instruction-dataset.texttext-generationn<1K0 likes39 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.