Hyukkyu/train-text2sql
Gretel Synthetic Text-to-SQL — Training, unified schema A seeded sample of gretelai/synthetic_text_to_sql, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources. Source gretelai/synthetic_text_to_sql @ 740ab236e645 Task question → schema and SQL Domain · languages code · eng Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-text2sql.
Gretel Synthetic Text-to-SQL — Training, unified schema
A seeded sample of `gretelai/synthetic_text_to_sql`, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources.
Schema
Files are Parquet, sorted by id, zstd-compressed, sharded at 500 MB. Every rule above is checked before publishing; provenance.json records the source revision, what changed, and the output file hashes.
What changed from the source
- sampled: a seeded random sample (seed 1) of up to 30,000 pairs
- reshaped: the question (
sql_prompt) is the query; the document is the schema (sql_context), a newline and the SQL (sql) - decontaminated (exact): a pair was dropped when its normalised query equals any evaluation query, or a positive equals a document of a test or dev corpus; a repeated query keeps its first pair
- decontaminated (near-duplicates): 0 passages and 0 queries that nearly copy a text of an evaluation set (word 13-grams for passages, 8-grams for queries; at least half shared with one text of the 23 test corpora (BEIR, RTEB, LitSearch) and the 3 dev corpora) were removed, and with them 0 queries in total
- text: leading and trailing whitespace stripped; otherwise as converted above
- ids re-keyed to
sha1(text)[:20]: 13 documents and 0 queries collapsed into identical texts - added a
titlecolumn filled with""(the source has none)
Hard negatives and teacher scores
Filled by the owner's annotation pipeline (annotation=jina35-u2) for the train split of the query set(s) below; queries without a labelled positive are left out.
- Candidates: dense retrieval with
jinaai/jina-embeddings-v5-text-smallover the full corpus to depth 1,000; 100 candidates per query drawn from the rank windows 1–30 (30), 31–100 (30), 101–300 (20), 301–1000 (20), the query's labelled positives excluded.rankis the dense rank;sourceisdensefor a mined row anddatasetfor a negative the source labels itself (those are kept for every query of the split, sampled or not). - Teacher:
jinaai/jina-reranker-v3.5, listwise: a query's positive and all of its candidates are scored together in one context of up to 32,768 tokens.scoreis the raw cosine score, one row per (query, positive) and per (query, candidate); a labelled negative that was also mined is scored once. Every document was cut to its first 1,024 reranker tokens before scoring (max_doc_tokens=1024). No filtering is applied to the tables.
from datasets import load_dataset
negatives = load_dataset("Hyukkyu/train-text2sql", "hard-negatives", split="train")
scores = load_dataset("Hyukkyu/train-text2sql", "teacher-scores", split="train")Jev judgments
judgments holds, for every query of the training sample (the queries with teacher scores), whether TypeSafe's Jev (jev-1.13.0) judged its training positive and its mined candidates relevant: p_yes is Jev's P(yes) for the source's question (e.g. does the passage answer the query?). They locate the false negatives among the mined candidates and the mislabelled positives.
- Requests. One request per query (
round0): its training positive, its candidates whose teacher score taken as (cos + 1) / 2 is at least 0.85 × the positive's (at most 12, the highest scores) and 4 random candidates below that, shuffled under neutral ids, one yes/no question per passage. Queries left with fewer than 10 candidates under their source's cutoff got their next hardest unjudged candidates in rounds 1–8 (8 per request), those still under 10 in rounds 9–10 (24 per request). The dataset's own labelled negatives were never sent. Texts were cut to 512 (query) and 512 (passage) tokens of thejina-embeddings-v5small tokenizer. Jev answers a request's passages in one context, so P(yes) is calibrated to these groups: the thresholds below apply to this table, not to single-pair calls. - Accuracy (an audit of 1237 pairs from the pilot's first requests (100 queries per source; the five long-query sources re-piloted at 512-token queries), labelled blind by an LLM (Claude), at the pilot's fixed thresholds 0.35 and 0.15): a candidate at P(yes) ≥ 0.35 was relevant 67% of the time inside the band (n = 350) and 45% below it (n = 87); one under 0.35 was relevant 6% (band, n = 387) and 1% (below the band, n = 210) of the time. A positive under 0.15 was mislabelled 100% of the time (n = 20) in the sources that keep the check; in dom-casehold, dom-clerc, dom-cornstack-py, dom-finqa10k, dom-gerlayqa, dom-investopedia, dom-lawse, dom-magicoder, dom-medmcqa, dom-pubmedqa, dom-s2orc, dom-tatqa, Jev's flags were right less often (0%–57% in this audit), under the 70% the check needs, so their positives are not checked.
- Use (the SPARSE loader,
annotation.filter.judge): a candidate at P(yes) ≥ its source's cutoff (below; fitted on 1,521 labelled pairs) is never a negative; a positive under 0.15 is replaced by the candidate Jev scores highest if that is ≥ 0.8, else the query is dropped; a candidate at ≥ 0.9 can become an extra positive. Comparep_yesas a float64 (it is stored as one).
Load it
from datasets import load_dataset
queries = load_dataset("Hyukkyu/train-text2sql", "queries", split="train")
corpus = load_dataset("Hyukkyu/train-text2sql", "corpus", split="train")
qrels = load_dataset("Hyukkyu/train-text2sql", "qrels", split="train")
negatives = load_dataset("Hyukkyu/train-text2sql", "hard-negatives", split="train")
scores = load_dataset("Hyukkyu/train-text2sql", "teacher-scores", split="train")
judgments = load_dataset("Hyukkyu/train-text2sql", "judgments", split="train")License and attribution
The data is redistributed under the source's terms — apache-2.0. All credit belongs to the original authors; see the source repository (https://huggingface.co/datasets/gretelai/synthetictextto_sql). This repository is an independent repackaging.
