Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CoIR-Retrieval /synthetic-text2sqlEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql.text100K<n<1M0 likes1.3k downloads2y agoHugging Face02sharkiefff /RBAC-Text2SQL-Benchmark RBAC-Text2SQL Benchmark Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural language question with a role policy; the model must either produce a correct SQL query that touches only authorized resources, or refuse with Sorry, I cannot answer. Code, evaluation harness, and reproduction instructions: https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.texttext-generation10K<n<100K1 likes180 downloads2mo agoHugging Face03CoIR-Retrieval /synthetic-text2sql-qrels Dataset Card for "synthetic-text2sql-qrels" More Information needed text100K<n<1M0 likes169 downloads2y agoHugging Face04CoIR-Retrieval /synthetic-text2sql-queries-corpusEmploying the CoIR evaluation framework's dataset version, utilize the code below for assessment: import coir from coir.data_loader import get_tasks from coir.evaluation import COIR from coir.models import YourCustomDEModel model_name = "intfloat/e5-base-v2" # Load the model model = YourCustomDEModel(model_name=model_name) # Get tasks #all task ["codetrans-dl","stackoverflow-qa","apps","codefeedback-mt","codefeedback-st","codetrans-contest","synthetic- # text2sql","cosqa","codesearchnet"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql-queries-corpus.text100K<n<1M2 likes161 downloads2y agoHugging Face05fahmiaziz /text2sql-dataset Dataset We built this dataset from several sources combining examples from: Wikisql Bird Spider Synthetic SQL samples This dataset has been cleaned and filtered by: Removing DDL/DML examples (INSERT, UPDATE, DELETE, etc.) De-duplicating examples based on hashing semantics of SQL and queries Filtering only SELECT-style analytical queries texttext-generation100K<n<1M1 likes158 downloads1y agoHugging Face06Hyukkyu /train-text2sql Gretel Synthetic Text-to-SQL — Training, unified schema A seeded sample of gretelai/synthetic_text_to_sql, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources. Source gretelai/synthetic_text_to_sql @ 740ab236e645 Task question → schema and SQL Domain · languages code · eng Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-text2sql.tabulartext-retrieval1M<n<10M0 likes137 downloads3d agoHugging Face07motherduckdb /duckdb-text2sql-25k Dataset Summary The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax. We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns. Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.text10K<n<100K43 likes133 downloads3y agoHugging Face08VikramPal /large-schema-text2sql-20k Large-Schema Text-to-SQL (20K) 20,020 text-to-SQL examples whose median database schema has 95 tables. Most text-to-SQL corpora hand the model a toy database. Spider averages about 5 tables per database; BIRD is in the same range. Real analytics work does not look like that — it looks like an ERP schema with 200 tables, 300 foreign keys, and eleven things called *_log, where the hard part is not writing the JOIN but finding the two tables worth joining. This dataset is that… See the full description on the dataset page: https://huggingface.co/datasets/VikramPal/large-schema-text2sql-20k.texttext-generation10K<n<100K0 likes103 downloads1mo agoHugging Face09shangrilar /ko_text2sqltext10K<n<100K19 likes93 downloads3y agoHugging Face10philikai /200k-Text2SQLtext100K<n<1M6 likes89 downloads3y agoHugging Face11lianghsun /bird-text2sql-bench Dataset Card for bird-text2sql-bench bird-text2sql-bench 是 BIRD(BIg Bench for Large-Scale Database Grounded Text-to-SQL) 官方訓練集之 OpenAI Messages 格式版本,共 9,428 筆。相較於 Spider 1.0,BIRD 使用真實大型資料庫(70 個,涵蓋電商、運動、教育、醫療等 37+ 領域),並提供 evidence(數值提示)欄位,本資料集將 evidence 以 ### Hint 段落併入 user prompt,形成可直接餵入 SFT pipeline 之 system / user / assistant 三 role 對話。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可同時作為 SFT 語料與 benchmark evaluation pipeline 之直接輸入。 Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/bird-text2sql-bench.texttext-generation1K<n<10K1 likes86 downloads6mo agoHugging Face12SuperMax991 /spider-text2sql SPIDER Text-to-SQL — Easy Access Version A clean, HuggingFace-native version of the SPIDER Text-to-SQL benchmark. The original SPIDER dataset requires manually downloading a ZIP file from the Spider website. This version makes it instantly accessible via load_dataset. What's Included Each row contains the question, gold SQL, the database identifier, and a pre-parsed compact schema string — everything needed to train or evaluate a Text-to-SQL model without any additional… See the full description on the dataset page: https://huggingface.co/datasets/SuperMax991/spider-text2sql.texttable-question-answering1K<n<10K2 likes85 downloads6mo agoHugging Face13johneze /chichewa-text2sql Chichewa Text-to-SQL The first structured Text-to-SQL benchmark for Chichewa, a low-resource Bantu language spoken by over 12 million people in Malawi and neighboring regions. The dataset contains 400 manually curated natural language–SQL pairs in both Chichewa (Nyanja) and English, grounded in a unified relational SQLite database covering five real-world domains from Malawi. Dataset Summary This benchmark was constructed to investigate the adaptation of Large… See the full description on the dataset page: https://huggingface.co/datasets/johneze/chichewa-text2sql.texttable-question-answeringn<1K2 likes77 downloads4mo agoHugging Face14eagle0504 /synthetic-text2sql-dataset Dataset Card for "synthetic-text2sql-dataset" Dataset Summary The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning. It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added: question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.textquestion-answering100K<n<1M1 likes59 downloads1y agoHugging Face15mohamed-ahmed-58059 /wikisql-text2sql WikiSQL Text-to-SQL (execution-ready) A cleaned, execution-ready repackaging of WikiSQL for text-to-SQL fine-tuning and execution-accuracy evaluation. Each example pairs a natural-language question with a gold SQL query over a single-table schema — and every split ships a real SQLite database so predicted SQL can be run and compared by result set (not string-matched). Splits Split Examples train 55,339 dev 8,263 test 15,519 Columns… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/wikisql-text2sql.texttext-generation10K<n<100K0 likes53 downloads4mo agoHugging Face16Healthy13 /Text2SQLtext100K<n<1M19 likes52 downloads3y agoHugging Face17InfiniFlow /text2sqltextn<1K13 likes52 downloads1y agoHugging Face18hoangphu7122002ai /text2sql_vitext100K<n<1M0 likes40 downloads3y agoHugging Face19chabab /text2sql-oracle-postgres Oracle / PostgreSQL text-to-SQL Instruction data for fine-tuning google/gemma-3-270m-it (or any chat model) to emit a single dialect-correct SQL statement. 804 rows, 402 Oracle / 402 PostgreSQL 7 schemas: hr, sales, banking, inventory, tickets, university, logistics Splits: 684 / 60 / 60 (grouped so paraphrases of the same SQL stay in one split) Load from datasets import load_dataset ds = load_dataset("chabab/text2sql-oracle-postgres") Record… See the full description on the dataset page: https://huggingface.co/datasets/chabab/text2sql-oracle-postgres.texttext-generationn<1K0 likes40 downloads2mo agoHugging Face20henrynguyen13 /text2sqltext100K<n<1M0 likes39 downloads2y agoHugging Face21Genies /text2sql-grpo-sql-r1-multi-steptext1K<n<10K1 likes38 downloads1y agoHugging Face22dipanjanS /text2sql-dataset Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/dipanjanS/text2sql-dataset.texttext-generation10K<n<100K0 likes34 downloads6mo agoHugging Face23lianghsun /spider-text2sql-bench Dataset Card for spider-text2sql-bench spider-text2sql-bench 是 Spider 1.0 官方訓練集之 OpenAI Messages 格式版本,共 7,000 筆,將原始之 question / schema / sql 重新組裝為 system / user / assistant 三 role 之對話結構。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可作為 Text-to-SQL 模型之 SFT 訓練語料,亦可直接用於 benchmark evaluation pipeline(以 user 作為 prompt,比對模型輸出與 assistant 之標準答案 SQL)。 Dataset Details Dataset Description Spider 1.0 為 Yale LILY Group 於 EMNLP 2018 發表之大規模跨領域 Text-to-SQL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/spider-text2sql-bench.texttext-generation1K<n<10K0 likes34 downloads6mo agoHugging Face24ManthanKulakarni /Text2SQLtext100K<n<1M0 likes32 downloads3y agoHugging Face25NafishZaldinanda /text2sql-omnisql-style Dialect: SQLite Dataset Source Paper Samples Used Notes Links Spider Spider: A Large-Scale Human-Labeled Dataset... 7,000 Seluruh training split digunakan. Link Google Drive Donwload BIRD23-Train-Filtered A BIg Bench for Large-Scale Database Grounded Text-to-SQLs 6,626 Menggunakan subset bird23-train-filtered. HuggingFace Dataset SynSQL-2.5M (Filtered) OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale 7,000 Hasil filtering berdasarkan question style dan… See the full description on the dataset page: https://huggingface.co/datasets/NafishZaldinanda/text2sql-omnisql-style.text10K<n<100K0 likes32 downloads4mo agoHugging Face26Shaleen123 /text2sql-v1text100K<n<1M0 likes31 downloads18d agoHugging Face27sanghyun89 /for_text2sql_wikiSQL_korean_by_google_translator_apitext100K<n<1M0 likes30 downloads2y agoHugging Face28fahmiaziz /text2sql-dataset-reasoningtext10K<n<100K1 likes30 downloads1y agoHugging Face29ssodha3 /text2sqldatatext1K<n<10K0 likes28 downloads3mo agoHugging Face30Chulinz /Text2SQL-Decisions Text2SQL-Decisions v0.3 28,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release. Configuration Rows Task License default 25,000 Four-candidate SQL plan selection on public sensor and e-commerce data CC BY 4.0 meter 3,081 Bounded meter planning decisions, including conversational context CC0-1.0 The configurations… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions.texttext-classification10K<n<100K0 likes27 downloads7h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.