datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bird-critic-1.0-sqlite
📢 Update 2026-03-23
We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. The schema file is included in the code repository https://github.com/bird-bench/BIRD-CRITIC-1/blob/main/baseline/data/sqlite_schema.jsonl
BIRD-CRITIC-1.0-SQLite
BIRD-Critic is the first SQL debugging… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-sqlite.livesqlbench-base-lite-sqlite
🚀 LiveSQLBench-Base-Lite
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks.
🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.six-gym-sqlite
📢 Update 2026-03-23
We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. This dataset is the train split of BIRD-Critic-SQLite, comprising 5,000 data instances for model training and development. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B.
📋 Dataset Structure
Below is a description of the dataset fields and additional… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-sqlite.bird-sqlite-sft-train
BIRD SQLite Text-to-SQL SFT Dataset
Supervised fine-tuning data for a SQLite-dialect text-to-SQL specialist model,
built from the BIRD benchmark train split.
Contents
7,483 train + 408 val examples spanning 69 distinct database schemas
(movie_platform, chicago_crime, hockey, mondial_geo, works_cycles, and 64
others), split by a stratified per-database 95/5 hold-out (sft_sqlite_ train.jsonl / sft_sqlite_val.jsonl) with zero exact overlap between them.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/hiimivantang/bird-sqlite-sft-train.Kaikki-Wiktionary-Ultimate-SQLite
Kaikki Ultimate Raw SQLite - World Dictionary Database (2026 Edition)
📌 Overview
This dataset is a high-performance SQLite conversion of the massive Kaikki.org (Wiktextract) raw data. It contains millions of lexical entries across thousands of languages, preserved in its absolute raw JSON format to ensure zero data loss.
This database is designed for developers, linguists, and AI researchers who need a structured, indexed, and offline-ready version of the world's most… See the full description on the dataset page: https://huggingface.co/datasets/wave101828228/Kaikki-Wiktionary-Ultimate-SQLite.zsql-sqlite-dpo
zsql-sqlite-dpo
This is a dataset for training machine learning models to convert natural
English language text into SQLite dialect SQL queries.
This dataset comprises 200,000 DPO pairs curated to support the rapid
development of text-to-SQL generation models. The uniqueness of this dataset
lies in its optimization process. The "chosen" field within each data pair
contains SQL queries that have been canonicalized, optimized, and which are
chosen from the candidate set which… See the full description on the dataset page: https://huggingface.co/datasets/zerolink/zsql-sqlite-dpo.gemma-4-e2b-SAE-sqlite
Gemma 4 E2B SAE SQLite Atlas
An exact, queryable SQLite representation of all 35 residual-stream
sparse autoencoders from
juiceb0xc0de/gemma-4-e2b-it-SAE.
The database contains 1,720,320 feature rows. Encoder and decoder
vectors preserve the source checkpoints' float32 values exactly.
Files
gemma-4-e2b-sae.sqlite3 — SQLite database (20.63 GiB)
manifest.json — source revision, dimensions, SHA-256, and integrity result
Database SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-SAE-sqlite.Gelbooru-SQLiteThis is db dump(s) of booru-typed databases.
The codebase (https://github.com/aria1th/Booru-Unified-Sqlite) will be used for creating DB, to handle various types of DB + allowing multiple DBs being loaded in same program.
Danbooru DB, mainly, will be updated at https://huggingface.co/datasets/KBlueLeaf/danbooru2023-sqlite too.
iraq-sqlite-dataqwen25-coder-7b-instruct_sqlitespider2_lite_sqlite_summerzCosmopedia_QA_RAG_JSON_SQLiteThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
🖥️ Demo Interface: Discord
Discord: https://discord.gg/Xe9tHFCS9h
**Custom RAG QA generation services can be made available for paying customers to process internal documentation. DM me on Discord if you are interested.Jeeney AI GPT Reloaded 207M/Cosmopedia Model Outputs Dataset
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Cosmopedia_QA_RAG_JSON_SQLite.sqli-test-httpsDanbooru2021-SQLite
Danbooru 2021 SQLite
Dataset Summary
This is the metadata of danbooru 2021 dataset in SQLite format.
https://gwern.net/danbooru2021
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation… See the full description on the dataset page: https://huggingface.co/datasets/cheryramneg/Danbooru2021-SQLite.qwen3_8b_sqlite_3xqwen3_8b_sqlite_3x-stepsSQLite_Training_Datasettext2sql-grpo-d6-e17-deema_sqlite-stepsqwen25-coder-7b-instruct_sqlite-stepssqli-test-http443text2sql-grpo-d6-e17-deema_sqlitetext2sql-sft-kumar-v9_sqlitetext2sql-sft-kumar-v9_sqlite-stepssqli-test-live-v2sqli-test-expensivesqli-test-divzerozcc-ir-sqlite-v1
zcc-ir-sqlite-v1
ZCC IR corpus harvested from SQLite 3.45 amalgamation. 2,054 functions compiled with ZCC and scored by PRIME Hamiltonian energy.
Stats
Total functions: 2054
LEGENDARY: 1485
EPIC: 569
Sources
sqlite: 1295
zcc_self: 754
curl: 5
Schema
Field
Type
Description
func_name
string
C function name
tier
string
LEGENDARY / EPIC / RARE / UNCOMMON / COMMON
tier_rank
int
Numeric tier (3=LEGENDARY)
node_count
int
IR node count… See the full description on the dataset page: https://huggingface.co/datasets/zkaedi/zcc-ir-sqlite-v1.HRMS_SQL_SQLitesqlite-0aa95099f5003dc99f599ab77ac0004950b281ef-reach-line-datasettext2sql-grpo-garry-v10_sqlite
