Chulinz/Text2SQL-Decisions-Benchmark
Text2SQL-Decisions Benchmark Seed v0.1 A small, frozen English evaluation seed, separate from Text2SQL-Decisions. It contains 24 SQL-choice questions on three newly authored synthetic schemas and 20 end-to-end questions for the Olist application. It is not a large or human-reviewed benchmark. Questions and gold were authored by the same Codex assistant before running Clef, so execution checks do not establish independent semantic review. SQL-choice track: 24 examples… See the full description on the dataset page: https://huggingface.co/datasets/Chulinz/Text2SQL-Decisions-Benchmark.
061
1{2 "training_rows": 25000,3 "exact_question_overlap": 0,4 "exact_state_overlap": 0,5 "new_choice_rows": 24,6 "gold_positions": {7 "plan_1": 6,8 "plan_2": 6,9 "plan_3": 6,10 "plan_4": 611 }12}13 