Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trl-lab /SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases Project page · Questions and SQL · Trained models · Python library · Citation This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.texttext-generation1K<n<10K1 likes279 downloads8d agoHugging Face02trl-lab /SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL Project page · Schemas and databases · Trained models · Python library · Citation SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.texttext-generation100K<n<1M1 likes212 downloads8d agoHugging Face03sharkiefff /RBAC-Text2SQL-Benchmark RBAC-Text2SQL Benchmark Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural language question with a role policy; the model must either produce a correct SQL query that touches only authorized resources, or refuse with Sorry, I cannot answer. Code, evaluation harness, and reproduction instructions: https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.texttext-generation10K<n<100K1 likes178 downloads2mo agoHugging Face04fahmiaziz /text2sql-dataset Dataset We built this dataset from several sources combining examples from: Wikisql Bird Spider Synthetic SQL samples This dataset has been cleaned and filtered by: Removing DDL/DML examples (INSERT, UPDATE, DELETE, etc.) De-duplicating examples based on hashing semantics of SQL and queries Filtering only SELECT-style analytical queries texttext-generation100K<n<1M1 likes159 downloads1y agoHugging Face05VikramPal /large-schema-text2sql-20k Large-Schema Text-to-SQL (20K) 20,020 text-to-SQL examples whose median database schema has 95 tables. Most text-to-SQL corpora hand the model a toy database. Spider averages about 5 tables per database; BIRD is in the same range. Real analytics work does not look like that — it looks like an ERP schema with 200 tables, 300 foreign keys, and eleven things called *_log, where the hard part is not writing the JOIN but finding the two tables worth joining. This dataset is that… See the full description on the dataset page: https://huggingface.co/datasets/VikramPal/large-schema-text2sql-20k.texttext-generation10K<n<100K0 likes102 downloads1mo agoHugging Face06lianghsun /bird-text2sql-bench Dataset Card for bird-text2sql-bench bird-text2sql-bench 是 BIRD(BIg Bench for Large-Scale Database Grounded Text-to-SQL) 官方訓練集之 OpenAI Messages 格式版本,共 9,428 筆。相較於 Spider 1.0,BIRD 使用真實大型資料庫(70 個,涵蓋電商、運動、教育、醫療等 37+ 領域),並提供 evidence(數值提示)欄位,本資料集將 evidence 以 ### Hint 段落併入 user prompt,形成可直接餵入 SFT pipeline 之 system / user / assistant 三 role 對話。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可同時作為 SFT 語料與 benchmark evaluation pipeline 之直接輸入。 Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/bird-text2sql-bench.texttext-generation1K<n<10K1 likes82 downloads6mo agoHugging Face07DanielRegaladoCardoso /text-to-sql-mix-v2 🔗 Part of the SQL Agent LLMOps project This dataset is one of three purpose-built training mixes for the SQL Agent LLMOps project — an end-to-end pipeline that converts natural-language questions into SQL, executes the query on user data, and renders a storytelling-grade visualization. Dataset Model trained Role 🤗 DanielRegaladoCardoso/text-to-sql-mix-v2 Qwen 2.5 Coder 7B NL question → SQL 🤗 DanielRegaladoCardoso/chart-reasoning-mix-v1Phi-3 Mini 3.8B (question +… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/text-to-sql-mix-v2.texttext-generation100K<n<1M0 likes79 downloads6mo agoHugging Face08Chulinz /Text2SQL-Decisions Text2SQL-Decisions v0.3 28,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release. Configuration Rows Task License default 25,000 Four-candidate SQL plan selection on public sensor and e-commerce data CC BY 4.0 meter 3,081 Bounded meter planning decisions, including conversational context CC0-1.0 The configurations… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions.texttext-classification10K<n<100K0 likes67 downloads1d agoHugging Face09Chulinz /Text2SQL-Decisions-Benchmark Text2SQL-Decisions Benchmark English evaluations for choosing SQL plans and answering database questions. This repository contains a 24-question synthetic test, 20 application-level questions, and model results on the separate Text2SQL-Decisions dataset. Model results Each model received the same question, database schema, and four SQL candidates. Reference answers were excluded from model input. Accuracy measures selection of the correct candidate. Model… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions-Benchmark.texttext-classificationn<1K0 likes61 downloads5h agoHugging Face10OpenDCAI /dataflow-demo-Text2SQL DataFlow Text-to-SQL Dataset This repository provides the dataset from the DataFlow project. It includes multiple JSON splits that showcase common Text-to-SQL training data formats, covering raw inputs, refined outputs, and augmented samples. The dataset released can be used to train and enhance large language models’ Text-to-SQL generation capabilities, improving their generalization performance on Text-to-SQL tasks. What’s in This Dataset Each split is a JSON array… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-Text2SQL.text-generation10K<n<100K1 likes59 downloads8mo agoHugging Face11rishhh /schemasage-sql-clean-text2sql SchemaSage-SQL Clean Text-to-SQL Dataset Dataset repo: rishhh/schemasage-sql-clean-text2sql This dataset contains normalized SchemaSage-SQL supervised examples with a consistent schema/question/answer format. Destructive SQL targets are converted to explicit refusal examples, invalid SQL targets are removed, and answer SQL that references tables or columns absent from the provided schema is filtered. Files text2sql_train.jsonl text2sql_validation.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/rishhh/schemasage-sql-clean-text2sql.text-generation0 likes56 downloads5mo agoHugging Face12eagle0504 /synthetic-text2sql-dataset Dataset Card for "synthetic-text2sql-dataset" Dataset Summary The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning. It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added: question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.textquestion-answering100K<n<1M1 likes51 downloads1y agoHugging Face13mohamed-ahmed-58059 /wikisql-text2sql WikiSQL Text-to-SQL (execution-ready) A cleaned, execution-ready repackaging of WikiSQL for text-to-SQL fine-tuning and execution-accuracy evaluation. Each example pairs a natural-language question with a gold SQL query over a single-table schema — and every split ships a real SQLite database so predicted SQL can be run and compared by result set (not string-matched). Splits Split Examples train 55,339 dev 8,263 test 15,519 Columns… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/wikisql-text2sql.texttext-generation10K<n<100K0 likes49 downloads4mo agoHugging Face14chabab /text2sql-oracle-postgres Oracle / PostgreSQL text-to-SQL Instruction data for fine-tuning google/gemma-3-270m-it (or any chat model) to emit a single dialect-correct SQL statement. 804 rows, 402 Oracle / 402 PostgreSQL 7 schemas: hr, sales, banking, inventory, tickets, university, logistics Splits: 684 / 60 / 60 (grouped so paraphrases of the same SQL stay in one split) Load from datasets import load_dataset ds = load_dataset("chabab/text2sql-oracle-postgres") Record… See the full description on the dataset page: https://huggingface.co/datasets/chabab/text2sql-oracle-postgres.texttext-generationn<1K0 likes41 downloads2mo agoHugging Face15dipanjanS /text2sql-dataset Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/dipanjanS/text2sql-dataset.texttext-generation10K<n<100K0 likes32 downloads6mo agoHugging Face16lianghsun /spider-text2sql-bench Dataset Card for spider-text2sql-bench spider-text2sql-bench 是 Spider 1.0 官方訓練集之 OpenAI Messages 格式版本,共 7,000 筆,將原始之 question / schema / sql 重新組裝為 system / user / assistant 三 role 之對話結構。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可作為 Text-to-SQL 模型之 SFT 訓練語料,亦可直接用於 benchmark evaluation pipeline(以 user 作為 prompt,比對模型輸出與 assistant 之標準答案 SQL)。 Dataset Details Dataset Description Spider 1.0 為 Yale LILY Group 於 EMNLP 2018 發表之大規模跨領域 Text-to-SQL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/spider-text2sql-bench.texttext-generation1K<n<10K0 likes32 downloads6mo agoHugging Face17chakshu2 /sample_synthetic_text_to_sql Sample Synthetic Text to SQL Dataset The dataset presents a substantial collection of expertly crafted Text-to-SQL samples, generated using open source LLM's and shared under an open-source license. Highlights of the dataset include: -- 1563 examples, divided into a training set of 1200 samples and a test set of 363 samples.-- Approximately less than 1 million tokens in total, with nearly 0.5 million representing code-specific tokens.-- Coverage spans a diverse range of 3… See the full description on the dataset page: https://huggingface.co/datasets/chakshu2/sample_synthetic_text_to_sql.question-answering1K<n<10K1 likes11 downloads2y agoHugging Face18mohamed-ahmed-58059 /text2sql-canonical-v3.1 text2sql-canonical-v3.1 Training data for a SQLite text-to-SQL model: 51,976 training rows and a 527-row validation split. Each row holds a database schema rendered as text, a natural-language question, an optional evidence hint, and the reference SQL. The mix caps synthetic data at 40 percent and gives real benchmark rows double weight. Split Rows BIRD Spider SynSQL train 51,976 33% 27% 40% val 527 34% 27% 40% Columns db_id, question, gold_sql… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/text2sql-canonical-v3.1.texttext-generation10K<n<100K0 likes10 downloads2mo agoHugging Face19StarsMakeGalaxy /enterprise-text2sql-curated-600 🚀 Enterprise Text-to-SQL & Analytical BI (Verbose CoT Reasoning) This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design. 📊 Dataset Composition & 3-Slice Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/enterprise-text2sql-curated-600.texttext-generationn<1K0 likes10 downloads2mo agoHugging Face20Brian-333 /text2sql-eval-resultstext-generation0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.