Team Ai
Datasetpublic

Chulinz/Text2SQL-Decisions

Text2SQL-Decisions v0.3 28,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release. Configuration Rows Task License default 25,000 Four-candidate SQL plan selection on public sensor and e-commerce data CC BY 4.0 meter 3,081 Bounded meter planning decisions, including conversational context CC0-1.0 The configurations… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions.

sourceHugging Facecc-by-4.0updated 1d agoView on Hugging Face
0likes67downloads
README.md161 linesDownload Raw Back to root
1---2pretty_name: Text2SQL-Decisions3license:4- cc-by-4.05- cc0-1.06language:7- en8task_categories:9- text-classification10- text-generation11tags:12- text-to-sql13- decision-model14- query-planning15- sqlite16- sensors17- ecommerce18- template-generated19size_categories:20- 10K<n<100K21configs:22- config_name: default23  data_files:24  - split: train25    path: data/train.jsonl26  - split: dev27    path: data/dev.jsonl28  - split: calibration29    path: data/calibration.jsonl30  - split: test31    path: data/test.jsonl32- config_name: meter33  data_files:34  - split: train35    path: meter/train.jsonl36  - split: dev37    path: meter/dev.jsonl38  - split: calibration39    path: meter/calibration.jsonl40  - split: test41    path: meter/test.jsonl42---43 44# Text2SQL-Decisions v0.345 4628,081 English-only examples across two separate configurations. Both are template-generated, execution-validated drafts, not human-reviewed benchmarks. No model has been fine-tuned as part of this release.47 48| Configuration | Rows | Task | License |49| --- | ---: | --- | --- |50| `default` | 25,000 | Four-candidate SQL plan selection on public sensor and e-commerce data | CC BY 4.0 |51| `meter` | 3,081 | Bounded meter planning decisions, including conversational context | CC0-1.0 |52 53The configurations have different supervision contracts. Load them explicitly and adapt each training builder; do not concatenate their rows without a contract-aware conversion. The default configuration is unchanged from v0.2.54 55## Synthetic meter configuration56 57Covers usage totals, rankings, period comparisons, daily anomalies, stale reporting, measured spike contributors, summaries, and conversational follow-ups. The synthetic fixture has 15 meters and 33,439 cumulative readings, a frozen Asia/Bangkok clock, and explicit missing/reset/invalid interval cases. Explanations identify measured contributors, not proven physical causes.58 59| Split | Rows |60| --- | ---: |61| train | 1,842 |62| dev | 445 |63| calibration | 408 |64| test | 386 |65 66All 3,081 gold decisions round-trip through the planner; 3,060 executable plans were checked against PostgreSQL and 21 require clarification. Splits keep 223 semantic families separate. Usage has an independent raw-reading oracle; other execution evidence is not independent human review. Templates and the shared fixture limit generalization claims.67 68```python69from datasets import load_dataset70meter = load_dataset("Chulinz/Text2SQL-Decisions", "meter")71```72 73Only `state` and `questions` are inputs; `decisions` supplies supervision. See [meter documentation](meter/README.md), [validation report](meter_validation_report.json), and [portable fixture instructions](meter_fixture/README.md). Meter examples and original synthetic fixture are CC0 under [meter/LICENSE](meter/LICENSE); the root LICENSE and UCI attribution below apply to the default configuration.74 75## Default configuration (unchanged v0.2)76 77**25,000 English-only text-to-SQL plan-selection examples**, grounded in public hydraulic sensor and e-commerce data. Every example has a question, schema, four candidate SQL plans, a gold choice, parameterized reference SQL and its result. The included SQLite database makes the queries executable.78 79This is an execution-validated, template-generated dataset draft. It is not a human-authored benchmark or evidence of production text-to-SQL performance. The initial bilingual v0.1 remains in repository history; this version replaces the main splits with English-only examples and a new grouped split assignment.80 81| Split | Examples | Purpose |82| --- | ---: | --- |83| train | 20,000 | Parameter fitting |84| dev | 2,000 | Development and model selection |85| calibration | 1,000 | Probability calibration after fitting |86| test | 2,000 | Held-out evaluation |87 88These are 25,000 distinct SQL/parameter pairs, not translated or paraphrased copies counted as additional examples. They still share deterministic grammar templates and underlying databases; row count is not the number of independently authored intents.89 90## Supported decisions91 92Questions cover COUNT, SUM, AVG, MIN and MAX; single and multiple AND/OR predicates; ISO date comparisons; actual NULL values; GROUP BY and HAVING; ranked groups with explicit tie-breaking; and a two-table invoice/line join, including distinct invoice counts versus line counts. Each split includes these major features. Exact per-source/category counts and integrity checks are in `validation_report.json`.93 94Candidate errors include wrong aggregation or metric, missing/reversed predicates, wrong AND/OR, invoice-versus-line counting, missing HAVING, and wrong grouping/order/limit. Candidates are unique and prepare in SQLite; their results differ from the gold result according to the independent Python evaluator. No artificial numeric answer offsets are used. Cases without enough such candidates are skipped. This selection favors distinguishable cases and is a limitation when evaluating ambiguity or empty results.95 96## Load and execute97 98```python99import json100import sqlite3101from datasets import load_dataset102from huggingface_hub import hf_hub_download103 104data = load_dataset("Chulinz/Text2SQL-Decisions")105row = data["train"][0]106state = json.loads(row["state"])107questions = json.loads(row["questions"])108gold = json.loads(row["gold"])109 110path = hf_hub_download("Chulinz/Text2SQL-Decisions", "database/source.sqlite", repo_type="dataset")111with sqlite3.connect(path) as connection:112    result = connection.execute(row["sql"], json.loads(row["sql_parameters"])).fetchall()113```114 115For decision-model training, only `state` and `questions` are model input; `gold` supplies supervision. Do not feed results, reference SQL, family IDs or other label-bearing metadata as input. The source schema itself can include observed outcome columns because the task is querying a database, not predicting those outcomes. Training-library compatibility must be checked with that library's dataset builder before a full run.116 117## Fields118 119All fields are strings, with structured values JSON-encoded where noted:120 121- `query`: English question.122- `state`: JSON object with question and relevant table schema.123- `questions`: JSON choice question, instructions and four candidate SQL/parameter plans.124- `gold`: JSON label and one-hot supervision. Confidence 1 describes the supervised target, not a calibrated prediction.125- `sql`, `sql_parameters`, `result`: reference SQL and JSON-encoded parameters/result.126- `id`, `family_id`, `split`: identity and structural partition.127- `source`, `domain`, `language`, `task`, `license`, `category`: provenance and task metadata. `category` is a slash-separated set of query-feature tags.128 129Input state is at most 2,600 UTF-8 bytes; state/questions fit a 32,000-byte budget. The format provides a four-way SQL-plan decision, not separately labeled table/column/filter questions for every planning stage.130 131## Source database132 133| Table | Rows | Meaning |134| --- | ---: | --- |135| `hydraulic_cycles` | 2,205 | Six measured sensor means per 60-second cycle and source component-condition annotations |136| `shopping_sessions` | 12,330 | Selected recorded-session columns including observed purchase outcome |137| `retail_lines` | 27,035 | Lines from eligible positive, non-cancelled 2-6-line invoices |138| `retail_invoices` | 7,019 | Source invoice country, date and nullable pseudonymous customer identifier |139 140The retail subset retains 620 invoices whose customer identifier is genuinely missing. No NULLs are fabricated. Excel dates are normalized to ISO YYYY-MM-DD, covering 2009-12-01 through 2011-12-09; time of day is discarded. Invoice keys retain a worksheet prefix to distinguish annual records. Customer identifiers are the public source's pseudonymous values, not names or contact details.141 142Retail invoices with invalid/nonpositive prices or quantities, fractional quantities or more than six lines were excluded in full. Prices are rounded HALF_UP to pennies. Sensor means are rounded to five decimal places and do not preserve waveform/frequency information. No claims of representativeness of all retail transactions or unseen machines are made. `sources.json` records source URLs, attribution, checksums and normalization details. TEP, Olist and BANKING77 are not included.143 144## Construction and evaluation limits145 146SQL templates produce the natural-language question and reference plan. Every gold query is executed and checked against an independent Python calculation, including NULL, empty-set, filtering, grouping, joins and tie-order semantics. Distractors prepare successfully and differ in expected result. Tests cover the compiler and independent evaluator. No LLM is called to generate questions or labels.147 148Structural families exclude literal values and remain wholly within one split. Selection balances available source/category groups within each assigned split and meets the stated quotas. The same source databases are shared across all splits. This is a query-family holdout, not an unseen-database, unseen-machine or chronological holdout. Repeated language templates and near-related query structures remain. Do not interpret execution validation as comprehensive human semantic review.149 150Unconstrained SQL generation, arbitrary joins, subqueries, window functions, free-form language diversity and deployed-application behavior are outside this release's validated scope. No model has been fine-tuned as part of dataset creation. Use per-source and per-category metrics, keep the test set out of tuning, and distinguish supplied-plan selection from end-to-end text-to-SQL accuracy.151 152## License and credit153 154The adapted dataset and documentation are distributed under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/), matching the license declared by UCI for all three sources. Attribute the original authors and this adaptation, link the license and describe further changes. See `LICENSE`; the original authors do not endorse this work.155 1561. Helwig, N., Pignanelli, E., and Schuetze, A. (2015). **Condition monitoring of hydraulic systems**. UCI Machine Learning Repository. [10.24432/C5CW21](https://doi.org/10.24432/C5CW21).1572. Chen, D. (2012). **Online Retail II**. UCI Machine Learning Repository. [10.24432/C5CG6D](https://doi.org/10.24432/C5CG6D).1583. Sakar, C., and Kastro, Y. (2018). **Online Shoppers Purchasing Intention Dataset**. UCI Machine Learning Repository. [10.24432/C5F88Q](https://doi.org/10.24432/C5F88Q).159 160Adaptation: Chulinz, **Text2SQL-Decisions**, 2026. Modifications include source filtering, normalization, sensor summaries, English SQL-question templates, candidate plans, exact supervision and grouped splits. Base-model licenses are separate.161