Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AsadIsmail /nl2sql-deduplicated NL2SQL Deduplicated Training Dataset A curated and deduplicated Text-to-SQL training dataset with 683,015 unique examples from 4 high-quality sources. 📊 Dataset Summary Total Examples: 683,015 unique question-SQL pairs Sources: Spider, SQaLe, Gretel Synthetic, SQL-Create-Context Deduplication Strategy: Input-only (question-based) with conflict resolution via quality priority Conflicts Resolved: 2,238 cases where same question had different SQL SQL Dialect: Standard SQL… See the full description on the dataset page: https://huggingface.co/datasets/AsadIsmail/nl2sql-deduplicated.texttext-generation1K<n<10K0 likes178 downloads10mo agoHugging Face02nl2sqlproj /spider2-nl2sql Dataset Details Dataset Description This dataset consists of data for the purpose of training a model to generate SQL code in response to a natural language prompt. The qa.csv table consists of these pairs, while the <dbms>_ddl.csv tables consist of the DDLs and sample data needed to verify the validity of generated SQL queries. Dataset Sources: Spider2 Repository Paper text-generation0 likes60 downloads1y agoHugging Face03zhangxiang666 /DS-NL2SQL DS-NL2SQL: A Benchmark for Dialect-Specific NL2SQL Paper: Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemCode Repository: weAIDB/Dial Dataset Overview Existing Text-to-SQL benchmarks (such as Spider and BIRD) predominantly focus on SQLite-compatible syntax, failing to capture the syntax specificity and heterogeneity inherent in real-world enterprise database dialects. To bridge this gap, we introduce DS-NL2SQL, a high-quality, multi-dialect NL2SQL benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zhangxiang666/DS-NL2SQL.table-question-answering1K<n<10K2 likes54 downloads7mo agoHugging Face04Shritama /nl2sqltext10K<n<100K2 likes45 downloads3y agoHugging Face05lorinma /NL2SQL_zh整合了3个中文数据集:追一科技NL2SQL,西湖大学的CSpider中文翻译,百度的DuSQL。 进行了大致的清洗,以及格式转换(alpaca): 假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。格式如下:{'sql':sql语句} MySQL数据库数据库结构如下:\n{表名(字段名...)}\n 其中:\n{表之间的主外键关联关系}\n 对于query:“{问题}”,给出相应的SQL语句,按照要求的格式返回,不进行任何解释。 其中,DuSQL最终结果是25004个。NL2SQL最终结果45919个,注意表名是乱码。CSpider,最终结果7786条,注意数据库是英文的,问题是中文的。 最终形成的文件,一共78706条,文件样例: { "instruction": "假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/NL2SQL_zh.text10K<n<100K19 likes30 downloads3y agoHugging Face06Rajpreet2206 /nl2sql-datasettext1K<n<10K1 likes27 downloads2y agoHugging Face07open-llm-leaderboard-old /details_uukuguy__speechless-nl2sql-ds-6.7b Dataset Card for Evaluation run of uukuguy/speechless-nl2sql-ds-6.7b Dataset automatically created during the evaluation run of model uukuguy/speechless-nl2sql-ds-6.7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-nl2sql-ds-6.7b.0 likes20 downloads3y agoHugging Face08riv25-aim410 /riv25_aim410_nl2sql_toolcalltext1K<n<10K0 likes19 downloads11mo agoHugging Face09AryaYT /southwest-dqe-nl2sql Southwest DQE NL2SQL Interview-oriented natural language to SQL examples for a Data Quality Engineer role centered on: Kafka → S3 Bronze (JSONL) → Glue/Spark → Silver/Gold Parquet → Redshift Contents Curated rows covering null profiling, window-function dedup, source/target reconciliation, SCD Type 2, quarantine reasons, and Bronze/Silver/Redshift balance checks Filtered slice of gretelai/synthetic_text_to_sql focused on analytics, joins, windows, CTEs, and… See the full description on the dataset page: https://huggingface.co/datasets/AryaYT/southwest-dqe-nl2sql.textn<1K0 likes18 downloads3mo agoHugging Face10DarianNLP /multilingual-nl2sql-datasets-gen_data_spidertabular10K<n<100K0 likes14 downloads9mo agoHugging Face11sirabhop /nl2sql_food_fieldnametextn<1K1 likes13 downloads3y agoHugging Face12JasperHaozhe /NL2SQL-Queriestext1K<n<10K0 likes12 downloads1y agoHugging Face13hajung /nl2sql_datasettextn<1K0 likes11 downloads3y agoHugging Face14selmoch /nl2sqltextn<1K0 likes10 downloads2y agoHugging Face15Lie24 /nl2sql-500ktext100K<n<1M1 likes10 downloads1y agoHugging Face16Arseniy-Sandalov /Med-NL2SQLtext10K<n<100K0 likes9 downloads2y agoHugging Face17DarianNLP /multilingual-nl2sql-datasets-filteredtabular1K<n<10K0 likes9 downloads9mo agoHugging Face18KyrieSun /nl2sql-patent-paper-100 NL2SQL-Patent-Paper-100 数据集 首个面向专利和论文检索的中文 NL2SQL 数据集,包含自然语言查询到结构化检索表达式的转换对。 数据集概述 数据集摘要 本数据集包含 100 条专利/论文检索领域的 NL2SQL 样本,每条数据包含: 中文自然语言查询(query) 对应的半结构化检索表达式(nl2sql) 英文翻译版本 这是更大规模数据集(742条)的样本子集,用于社区验证和基线测试。 支持的任务 Text-to-SQL:自然语言到结构化查询转换 语义解析:自然语言理解到形式化表达 信息检索:专利/论文检索查询理解 机器翻译:中英技术文本翻译 语言 中文(简体) 英文 领域 专利检索 学术论文检索 知识产权 数据集结构 数据字段 字段名 类型 描述 示例 id int 唯一标识符 4… See the full description on the dataset page: https://huggingface.co/datasets/KyrieSun/nl2sql-patent-paper-100.textn<1K1 likes9 downloads6mo agoHugging Face19jerichosiahaya /nl2sql_qna_pairtext10K<n<100K0 likes8 downloads9mo agoHugging Face20Dortp58 /nl2sql_datasettext1K<n<10K0 likes7 downloads1y agoHugging Face21TieLu /NL2SQL_zh整合了3个中文数据集:追一科技NL2SQL,西湖大学的CSpider中文翻译,百度的DuSQL。 进行了大致的清洗,以及格式转换(alpaca): 假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。格式如下:{'sql':sql语句} MySQL数据库数据库结构如下:\n{表名(字段名...)}\n 其中:\n{表之间的主外键关联关系}\n 对于query:“{问题}”,给出相应的SQL语句,按照要求的格式返回,不进行任何解释。 其中,DuSQL最终结果是25004个。NL2SQL最终结果45919个,注意表名是乱码。CSpider,最终结果7786条,注意数据库是英文的,问题是中文的。 最终形成的文件,一共78706条,文件样例: { "instruction": "假设你是一个数据库SQL专家,下面我会给出一个MySQL数据库的信息,请根据问题,帮我生成相应的SQL语句。当前时间为2023年。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/TieLu/NL2SQL_zh.text10K<n<100K0 likes7 downloads5mo agoHugging Face22ChengSong /nl2sql_general_ability Dataset Card for "nl2sql_general_ability" More Information needed text1K<n<10K0 likes6 downloads3y agoHugging Face23NormalMatt /nl2sqltext1K<n<10K0 likes6 downloads2y agoHugging Face24DarianNLP /multilingual-nl2sql-datasets-gen_datatext10K<n<100K0 likes6 downloads9mo agoHugging Face25hajung /nl2sqldatatextn<1K0 likes5 downloads3y agoHugging Face26ManoharPalanisamy /NL2SQLtextn<1K0 likes5 downloads2y agoHugging Face27simone-papicchio /nl2sql-reasoning-tracegatedtext1K<n<10K0 likes5 downloads2y agoHugging Face28symbolzh /nl2sql-1229-10ktabular10K<n<100K0 likes5 downloads9mo agoHugging Face29replysadiq /qatar-customs-nl2sql Qatar Customs NL2SQL — Arabic Natural Language to SAP HANA SQL Fine-tuning dataset for converting Arabic natural language questions into SAP HANA SQL queries for the Qatar General Directorate of Customs (GDC) database. Dataset Overview Metric Value Total question sets 99 Training examples 320 (3 fuzz levels × 80 questions + 80 follow-ups) Test examples 76 (3 fuzz levels × 19 questions + 19 follow-ups) Language Arabic 🇶🇦 (formal + dialectal) SQL… See the full description on the dataset page: https://huggingface.co/datasets/replysadiq/qatar-customs-nl2sql.textn<1K0 likes5 downloads6mo agoHugging Face30ChengSong /nl2sql_general_ability_enhanced Dataset Card for "nl2sql_general_ability_enhanced" More Information needed text100K<n<1M2 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.