Team Ai
Datasetpublic

Chulinz/Text2SQL-Decisions-Benchmark

Text2SQL-Decisions Benchmark Seed v0.1 A small, frozen English evaluation seed, separate from Text2SQL-Decisions. It contains 24 SQL-choice questions on three newly authored synthetic schemas and 20 end-to-end questions for the Olist application. It is not a large or human-reviewed benchmark. Questions and gold were authored by the same Codex assistant before running Clef, so execution checks do not establish independent semantic review. SQL-choice track: 24 examples… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions-Benchmark.

sourceHugging Facecc0-1.0updated 14h agoView on Hugging Face
0likes
Dataset Card

Text2SQL-Decisions Benchmark Seed v0.1

A small, frozen English evaluation seed, separate from Text2SQL-Decisions. It contains 24 SQL-choice questions on three newly authored synthetic schemas and 20 end-to-end questions for the Olist application. It is not a large or human-reviewed benchmark. Questions and gold were authored by the same Codex assistant before running Clef, so execution checks do not establish independent semantic review.

SQL-choice track: 24 examples

Eight cases each cover purchases/refunds, inventory, and subscriptions. Topics include NULLs, conditional aggregation, anti-joins, refund fan-out, grouped HAVING, explicit ordering, date boundaries and a window function. fixture.sql recreates the SQLite database using synthetic rows only. Every gold result was checked against a predeclared expected result; all four candidates execute and return distinct results. Positions of gold choices are balanced, six per position.

Only state and questions are model input; gold, sql, result and metadata are labels or evaluation artifacts. The model selects one of four supplied SQL queries. No free-form SQL generation is measured. The schemas and exact questions are absent from the 25k training-dataset release; simple SQL operations still overlap, and this is not a guarantee against model pretraining contamination.

python
from datasets import load_dataset
choice = load_dataset("Chulinz/Text2SQL-Decisions-Benchmark", "default", split="test")
ecommerce = load_dataset("Chulinz/Text2SQL-Decisions-Benchmark", "ecommerce_e2e", split="test")

Cloudflare Clef Flash scored 24/24 with Cloudflare-only routing. This tiny result is consistent with an easy seed and should not be presented as proof of broad text-to-SQL ability. See baseline.json for per-case outputs and the frozen data checksum. No prompt or harness changes were made in response to these results.

End-to-end track: 20 examples

These questions are sent to the actual Olist harness, without supplying SQL candidates. The harness chooses its plan, executes PostgreSQL and returns an answer or alternatives. Eighteen questions have executable raw-source gold SQL; two expect refusal. Question and SQL fields are newly authored. Olist data are not included and must be obtained separately under their own CC BY-NC-SA 4.0 terms; this repository's CC0 dedication does not relicense Olist.

Use decisionmodel-text-to-sql at commit 71aa13af7af07d48ce489b1bd062c3c98ec26a97, complete its database setup, then run:

sh
npm run test:live -- --cases /absolute/path/ecommerce-e2e.json

Frozen-run results: 12 accepted correct, 4 accepted wrong against gold, 2 offered a choice containing the correct answer, 2 refused as expected, 0 API errors. Direct correct answers on answerable cases are 12/18 (66.67%); choices containing the correct answer are reported separately, not counted as direct success. The scorer compares result values and enforces ordering where requested. See ecommerce-e2e-results.json.

The delivery-date case's gold compares timestamps without requiring order_status='delivered'; the harness adds that status. This is a potential interpretation dispute requiring human review. The frozen gold is retained and the mismatch is disclosed, rather than changing labels after seeing results. Other observed mismatches include dropped zero/NULL filters and incorrect requested ordering. This seed does not yet cover conversational chains or all eight meter use cases.

Reproducibility and scope

manifest.json records checksums frozen before model evaluation. overlap-audit.json checks exact question/state overlap with all 25k source rows. build-benchmark.py supplies the synthetic fixture, expected results and deterministic candidate rotation; it makes no API calls. Keep this repository out of fine-tuning data, and version any later revisions. Public exposure means future contamination remains possible.

The existing 2,000-row test split is a separate evaluation: Clef scored 1,967/2,000 (98.35%) on that split in the full rerun. It uses supplied SQL candidates and shared training-source databases, so it is not comparable to the end-to-end rate above. training-dataset-test-metrics.json preserves that result. Train/dev/calibration audit scores are not held-out benchmark scores.

Newly authored questions, SQL, synthetic fixture and documentation are dedicated under CC0-1.0. Model/provider and external source-data terms remain separate. No fine-tuning was performed.

Complete 25k dataset audit

All 25,000 source rows were evaluated with the same Cloudflare-only decision request contract. See training-dataset-full-audit.json for separate train/dev/calibration/test scores. The held-out test remains 1,967/2,000 (98.35%); the all-split total is not a held-out benchmark. Successful results from an initial rate-limited run were reused only when request hashes matched. API-reported costs exclude any unreported cost for rate-limited attempts.