Chulinz/Text2SQL-Decisions
Text2SQL-Decisions v0.3 28,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release. Configuration Rows Task License default 25,000 Four-candidate SQL plan selection on public sensor and e-commerce data CC BY 4.0 meter 3,081 Bounded meter planning decisions, including conversational context CC0-1.0 The configurations… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions.
Text2SQL-Decisions v0.3
28,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release.
The configurations have different supervision contracts. Load them explicitly and adapt each training builder; do not concatenate their rows without a contract-aware conversion. The default configuration is unchanged from v0.2.
Synthetic meter configuration
Covers usage totals, rankings, period comparisons, daily anomalies, stale reporting, measured spike contributors, summaries, and conversational follow-ups. The synthetic fixture has 15 meters and 33,439 cumulative readings, a frozen Asia/Bangkok clock, and explicit missing/reset/invalid interval cases. Explanations identify measured contributors, not proven physical causes.
All 3,081 gold decisions round-trip through the planner; 3,060 executable plans were checked against PostgreSQL and 21 require clarification. Splits keep 223 semantic families separate. Usage has an independent raw-reading oracle; other execution evidence is not independent human review. Templates and the shared fixture limit generalization claims.
from datasets import load_dataset
meter = load_dataset("Chulinz/Text2SQL-Decisions", "meter")Only state and questions are inputs; decisions supplies supervision. See meter documentation, validation report, and portable fixture instructions. Meter examples and original synthetic fixture are CC0 under meter/LICENSE; the root LICENSE and UCI attribution below apply to the default configuration.
Default configuration (unchanged v0.2)
25,000 English-only text-to-SQL plan-selection examples, grounded in public hydraulic sensor and e-commerce data. Every example has a question, schema, four candidate SQL plans, a gold choice, parameterized reference SQL and its result. The included SQLite database makes the queries executable.
This is an execution-validated, template-generated dataset draft. It is not a human-authored benchmark or evidence of production text-to-SQL performance. The initial bilingual v0.1 remains in repository history; this version replaces the main splits with English-only examples and a new grouped split assignment.
These are 25,000 distinct SQL/parameter pairs, not translated or paraphrased copies counted as additional examples. They still share deterministic grammar templates and underlying databases; row count is not the number of independently authored intents.
Supported decisions
Questions cover COUNT, SUM, AVG, MIN and MAX; single and multiple AND/OR predicates; ISO date comparisons; actual NULL values; GROUP BY and HAVING; ranked groups with explicit tie-breaking; and a two-table invoice/line join, including distinct invoice counts versus line counts. Each split includes these major features. Exact per-source/category counts and integrity checks are in validation_report.json.
Candidate errors include wrong aggregation or metric, missing/reversed predicates, wrong AND/OR, invoice-versus-line counting, missing HAVING, and wrong grouping/order/limit. Candidates are unique and prepare in SQLite; their results differ from the gold result according to the independent Python evaluator. No artificial numeric answer offsets are used. Cases without enough such candidates are skipped. This selection favors distinguishable cases and is a limitation when evaluating ambiguity or empty results.
Load and execute
import json
import sqlite3
from datasets import load_dataset
from huggingface_hub import hf_hub_download
data = load_dataset("Chulinz/Text2SQL-Decisions")
row = data["train"][0]
state = json.loads(row["state"])
questions = json.loads(row["questions"])
gold = json.loads(row["gold"])
path = hf_hub_download("Chulinz/Text2SQL-Decisions", "database/source.sqlite", repo_type="dataset")
with sqlite3.connect(path) as connection:
result = connection.execute(row["sql"], json.loads(row["sql_parameters"])).fetchall()For decision-model training, only state and questions are model input; gold supplies supervision. Do not feed results, reference SQL, family IDs or other label-bearing metadata as input. The source schema itself can include observed outcome columns because the task is querying a database, not predicting those outcomes. Training-library compatibility must be checked with that library's dataset builder before a full run.
Fields
All fields are strings, with structured values JSON-encoded where noted:
query: English question.state: JSON object with question and relevant table schema.questions: JSON choice question, instructions and four candidate SQL/parameter plans.gold: JSON label and one-hot supervision. Confidence 1 describes the supervised target, not a calibrated prediction.sql,sql_parameters,result: reference SQL and JSON-encoded parameters/result.id,family_id,split: identity and structural partition.source,domain,language,task,license,category: provenance and task metadata.categoryis a slash-separated set of query-feature tags.
Input state is at most 2,600 UTF-8 bytes; state/questions fit a 32,000-byte budget. The format provides a four-way SQL-plan decision, not separately labeled table/column/filter questions for every planning stage.
Source database
The retail subset retains 620 invoices whose customer identifier is genuinely missing. No NULLs are fabricated. Excel dates are normalized to ISO YYYY-MM-DD, covering 2009-12-01 through 2011-12-09; time of day is discarded. Invoice keys retain a worksheet prefix to distinguish annual records. Customer identifiers are the public source's pseudonymous values, not names or contact details.
Retail invoices with invalid/nonpositive prices or quantities, fractional quantities or more than six lines were excluded in full. Prices are rounded HALF_UP to pennies. Sensor means are rounded to five decimal places and do not preserve waveform/frequency information. No claims of representativeness of all retail transactions or unseen machines are made. sources.json records source URLs, attribution, checksums and normalization details. TEP, Olist and BANKING77 are not included.
Construction and evaluation limits
SQL templates produce the natural-language question and reference plan. Every gold query is executed and checked against an independent Python calculation, including NULL, empty-set, filtering, grouping, joins and tie-order semantics. Distractors prepare successfully and differ in expected result. Tests cover the compiler and independent evaluator. No LLM is called to generate questions or labels.
Structural families exclude literal values and remain wholly within one split. Selection balances available source/category groups within each assigned split and meets the stated quotas. The same source databases are shared across all splits. This is a query-family holdout, not an unseen-database, unseen-machine or chronological holdout. Repeated language templates and near-related query structures remain. Do not interpret execution validation as comprehensive human semantic review.
Unconstrained SQL generation, arbitrary joins, subqueries, window functions, free-form language diversity and deployed-application behavior are outside this release's validated scope. No model has been fine-tuned as part of dataset creation. Use per-source and per-category metrics, keep the test set out of tuning, and distinguish supplied-plan selection from end-to-end text-to-SQL accuracy.
License and credit
The adapted dataset and documentation are distributed under CC BY 4.0, matching the license declared by UCI for all three sources. Attribute the original authors and this adaptation, link the license and describe further changes. See LICENSE; the original authors do not endorse this work.
- Helwig, N., Pignanelli, E., and Schuetze, A. (2015). Condition monitoring of hydraulic systems. UCI Machine Learning Repository. 10.24432/C5CW21.
- Chen, D. (2012). Online Retail II. UCI Machine Learning Repository. 10.24432/C5CG6D.
- Sakar, C., and Kastro, Y. (2018). Online Shoppers Purchasing Intention Dataset. UCI Machine Learning Repository. 10.24432/C5F88Q.
Adaptation: Chulinz, Text2SQL-Decisions, 2026. Modifications include source filtering, normalization, sensor summaries, English SQL-question templates, candidate plans, exact supervision and grouped splits. Base-model licenses are separate.
