Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AsadIsmail /nl2sql-deduplicated NL2SQL Deduplicated Training Dataset A curated and deduplicated Text-to-SQL training dataset with 683,015 unique examples from 4 high-quality sources. 📊 Dataset Summary Total Examples: 683,015 unique question-SQL pairs Sources: Spider, SQaLe, Gretel Synthetic, SQL-Create-Context Deduplication Strategy: Input-only (question-based) with conflict resolution via quality priority Conflicts Resolved: 2,238 cases where same question had different SQL SQL Dialect: Standard SQL… See the full description on the dataset page: https://huggingface.co/datasets/AsadIsmail/nl2sql-deduplicated.texttext-generation1K<n<10K0 likes156 downloads10mo agoHugging Face02nl2sqlproj /spider2-nl2sql Dataset Details Dataset Description This dataset consists of data for the purpose of training a model to generate SQL code in response to a natural language prompt. The qa.csv table consists of these pairs, while the <dbms>_ddl.csv tables consist of the DDLs and sample data needed to verify the validity of generated SQL queries. Dataset Sources: Spider2 Repository Paper text-generation0 likes56 downloads1y agoHugging Face03zhangxiang666 /DS-NL2SQL DS-NL2SQL: A Benchmark for Dialect-Specific NL2SQL Paper: Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemCode Repository: weAIDB/Dial Dataset Overview Existing Text-to-SQL benchmarks (such as Spider and BIRD) predominantly focus on SQLite-compatible syntax, failing to capture the syntax specificity and heterogeneity inherent in real-world enterprise database dialects. To bridge this gap, we introduce DS-NL2SQL, a high-quality, multi-dialect NL2SQL benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zhangxiang666/DS-NL2SQL.table-question-answering1K<n<10K2 likes50 downloads7mo agoHugging Face04Shritama /nl2sqltext10K<n<100K2 likes46 downloads3y agoHugging Face05lorinma /NL2SQL_zh整合了3个中文数据集:追一科技NL2SQL,西湖大学的CSpider中文翻译,百度的DuSQL。 进行了大致的清洗,以及格式转换(alpaca): 假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。格式如下:{'sql':sql语句} MySQL数据库数据库结构如下:\n{表名(字段名...)}\n 其中:\n{表之间的主外键关联关系}\n 对于query:“{问题}”,给出相应的SQL语句,按照要求的格式返回,不进行任何解释。 其中,DuSQL最终结果是25004个。NL2SQL最终结果45919个,注意表名是乱码。CSpider,最终结果7786条,注意数据库是英文的,问题是中文的。 最终形成的文件,一共78706条,文件样例: { "instruction": "假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/NL2SQL_zh.text10K<n<100K19 likes31 downloads3y agoHugging Face06Rajpreet2206 /nl2sql-datasettext1K<n<10K1 likes25 downloads2y agoHugging Face07open-llm-leaderboard-old /details_uukuguy__speechless-nl2sql-ds-6.7b Dataset Card for Evaluation run of uukuguy/speechless-nl2sql-ds-6.7b Dataset automatically created during the evaluation run of model uukuguy/speechless-nl2sql-ds-6.7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-nl2sql-ds-6.7b.0 likes20 downloads3y agoHugging Face08AryaYT /southwest-dqe-nl2sql Southwest DQE NL2SQL Interview-oriented natural language to SQL examples for a Data Quality Engineer role centered on: Kafka → S3 Bronze (JSONL) → Glue/Spark → Silver/Gold Parquet → Redshift Contents Curated rows covering null profiling, window-function dedup, source/target reconciliation, SCD Type 2, quarantine reasons, and Bronze/Silver/Redshift balance checks Filtered slice of gretelai/synthetic_text_to_sql focused on analytics, joins, windows, CTEs, and… See the full description on the dataset page: https://huggingface.co/datasets/AryaYT/southwest-dqe-nl2sql.textn<1K0 likes20 downloads2mo agoHugging Face09riv25-aim410 /riv25_aim410_nl2sql_toolcalltext1K<n<10K0 likes18 downloads11mo agoHugging Face10DarianNLP /multilingual-nl2sql-datasets-gen_data_spidertabular10K<n<100K0 likes17 downloads9mo agoHugging Face11JasperHaozhe /NL2SQL-Queriestext1K<n<10K0 likes15 downloads1y agoHugging Face12sirabhop /nl2sql_food_fieldnametextn<1K1 likes13 downloads3y agoHugging Face13Lie24 /nl2sql-500ktext100K<n<1M1 likes13 downloads1y agoHugging Face14DarianNLP /multilingual-nl2sql-datasets-filteredtabular1K<n<10K0 likes12 downloads9mo agoHugging Face15hajung /nl2sql_datasettextn<1K0 likes11 downloads3y agoHugging Face16selmoch /nl2sqltextn<1K0 likes10 downloads2y agoHugging Face17Dortp58 /nl2sql_datasettext1K<n<10K0 likes10 downloads1y agoHugging Face18KyrieSun /nl2sql-patent-paper-100 NL2SQL-Patent-Paper-100 数据集 首个面向专利和论文检索的中文 NL2SQL 数据集,包含自然语言查询到结构化检索表达式的转换对。 数据集概述 数据集摘要 本数据集包含 100 条专利/论文检索领域的 NL2SQL 样本,每条数据包含: 中文自然语言查询(query) 对应的半结构化检索表达式(nl2sql) 英文翻译版本 这是更大规模数据集(742条)的样本子集,用于社区验证和基线测试。 支持的任务 Text-to-SQL:自然语言到结构化查询转换 语义解析:自然语言理解到形式化表达 信息检索:专利/论文检索查询理解 机器翻译:中英技术文本翻译 语言 中文(简体) 英文 领域 专利检索 学术论文检索 知识产权 数据集结构 数据字段 字段名 类型 描述 示例 id int 唯一标识符 4… See the full description on the dataset page: https://huggingface.co/datasets/KyrieSun/nl2sql-patent-paper-100.textn<1K1 likes10 downloads6mo agoHugging Face19ChengSong /nl2sql_general_ability Dataset Card for "nl2sql_general_ability" More Information needed text1K<n<10K0 likes8 downloads3y agoHugging Face20Arseniy-Sandalov /Med-NL2SQLtext10K<n<100K0 likes8 downloads2y agoHugging Face21DarianNLP /multilingual-nl2sql-datasets-gen_datatext10K<n<100K0 likes8 downloads9mo agoHugging Face22simone-papicchio /nl2sql-reasoning-tracegatedtext1K<n<10K0 likes7 downloads2y agoHugging Face23jerichosiahaya /nl2sql_qna_pairtext10K<n<100K0 likes7 downloads9mo agoHugging Face24TieLu /NL2SQL_zh整合了3个中文数据集:追一科技NL2SQL,西湖大学的CSpider中文翻译,百度的DuSQL。 进行了大致的清洗,以及格式转换(alpaca): 假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。格式如下:{'sql':sql语句} MySQL数据库数据库结构如下:\n{表名(字段名...)}\n 其中:\n{表之间的主外键关联关系}\n 对于query:“{问题}”,给出相应的SQL语句,按照要求的格式返回,不进行任何解释。 其中,DuSQL最终结果是25004个。NL2SQL最终结果45919个,注意表名是乱码。CSpider,最终结果7786条,注意数据库是英文的,问题是中文的。 最终形成的文件,一共78706条,文件样例: { "instruction": "假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/TieLu/NL2SQL_zh.text10K<n<100K0 likes7 downloads5mo agoHugging Face25ChengSong /nl2sql_general_ability_enhanced Dataset Card for "nl2sql_general_ability_enhanced" More Information needed text100K<n<1M2 likes6 downloads3y agoHugging Face26symbolzh /nl2sql-1229-10ktabular10K<n<100K0 likes6 downloads9mo agoHugging Face27hajung /nl2sqldatatextn<1K0 likes5 downloads3y agoHugging Face28ManoharPalanisamy /NL2SQLtextn<1K0 likes5 downloads2y agoHugging Face29replysadiq /qatar-customs-nl2sql Qatar Customs NL2SQL — Arabic Natural Language to SAP HANA SQL Fine-tuning dataset for converting Arabic natural language questions into SAP HANA SQL queries for the Qatar General Directorate of Customs (GDC) database. Dataset Overview Metric Value Total question sets 99 Training examples 320 (3 fuzz levels × 80 questions + 80 follow-ups) Test examples 76 (3 fuzz levels × 19 questions + 19 follow-ups) Language Arabic 🇶🇦 (formal + dialectal) SQL… See the full description on the dataset page: https://huggingface.co/datasets/replysadiq/qatar-customs-nl2sql.textn<1K0 likes5 downloads5mo agoHugging Face30JasperHaozhe /NL2SQL-Database0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.