Team Ai
Datasetpublic

Hyukkyu/train-text2sql

Gretel Synthetic Text-to-SQL — Training, unified schema A seeded sample of gretelai/synthetic_text_to_sql, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources. Source gretelai/synthetic_text_to_sql @ 740ab236e645 Task question → schema and SQL Domain · languages code · eng Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-text2sql.

sourceHugging Faceapache-2.0updated 14h agoView on Hugging Face
0likes117downloads
Dataset Card

Gretel Synthetic Text-to-SQL — Training, unified schema

A seeded sample of `gretelai/synthetic_text_to_sql`, made into retrieval training pairs and reshaped into the strict schema shared by every dataset in this collection. One of the 15 domain sources (code, medical, science, finance, legal) added to the collection's general sources.

Source`gretelai/synthetic_text_to_sql` @ 740ab236e645
Taskquestion → schema and SQL
Domain · languagescode · eng
Queries / documents / qrels30,000 / 29,987 / 30,000
Qrels per querymin 1 · mean 1.0 · max 1
Score values2 ×30,000 (2: the first positive, 1: any other)
Layoutqueries · corpus · qrels · hard-negatives · teacher-scores, split train
Splitscorpus: train · hard-negatives: train · judgments: train · qrels: train · queries: train · teacher-scores: train
Hard negativessources: dense · 2,970,047 rows
Teacher scoresjinaai/jina-reranker-v3.5 · 3,000,047 rows (positives included)
Judgmentsjudgments: typesafe/jev-1.13.0 · 399,530 rows
Idssha1(text)[:20]; identical texts collapse to one document / query
Licenseapache-2.0

Schema

configcolumnsrules
queriesid: string, text: stringids unique and non-empty; every query has ≥ 1 qrel
corpusid: string, title: string, text: stringtitle is always present ("" when the source has none)
qrelsquery-id: string, corpus-id: string, score: int32referential integrity to both tables; no duplicate pairs; no floats
hard-negativesquery-id: string, corpus-id: string, rank: int32, source: stringone row per negative; (query-id, corpus-id, source) unique; never a labelled positive of the same query
teacher-scoresquery-id: string, corpus-id: string, teacher: string, score: float32one row per scored pair (positives included); a row means scored — never a placeholder
judgmentsquery-id: string, corpus-id: string, judge: string, role: string, p_yes: float64, round: int32one row per judged pair; role is positive (the training positive) or candidate (a mined candidate, never a labelled positive or a labelled negative); p_yes in [0, 1]; round 0 the first request, 1.. the top-ups

Files are Parquet, sorted by id, zstd-compressed, sharded at 500 MB. Every rule above is checked before publishing; provenance.json records the source revision, what changed, and the output file hashes.

What changed from the source

  • —sampled: a seeded random sample (seed 1) of up to 30,000 pairs
  • —reshaped: the question (sql_prompt) is the query; the document is the schema (sql_context), a newline and the SQL (sql)
  • —decontaminated (exact): a pair was dropped when its normalised query equals any evaluation query, or a positive equals a document of a test or dev corpus; a repeated query keeps its first pair
  • —decontaminated (near-duplicates): 0 passages and 0 queries that nearly copy a text of an evaluation set (word 13-grams for passages, 8-grams for queries; at least half shared with one text of the 23 test corpora (BEIR, RTEB, LitSearch) and the 3 dev corpora) were removed, and with them 0 queries in total
  • —text: leading and trailing whitespace stripped; otherwise as converted above
  • —ids re-keyed to sha1(text)[:20]: 13 documents and 0 queries collapsed into identical texts
  • —added a title column filled with "" (the source has none)

Hard negatives and teacher scores

Filled by the owner's annotation pipeline (annotation=jina35-u2) for the train split of the query set(s) below; queries without a labelled positive are left out.

  • —Candidates: dense retrieval with jinaai/jina-embeddings-v5-text-small over the full corpus to depth 1,000; 100 candidates per query drawn from the rank windows 1–30 (30), 31–100 (30), 101–300 (20), 301–1000 (20), the query's labelled positives excluded. rank is the dense rank; source is dense for a mined row and dataset for a negative the source labels itself (those are kept for every query of the split, sampled or not).
  • —Teacher: jinaai/jina-reranker-v3.5, listwise: a query's positive and all of its candidates are scored together in one context of up to 32,768 tokens. score is the raw cosine score, one row per (query, positive) and per (query, candidate); a labelled negative that was also mined is scored once. Every document was cut to its first 1,024 reranker tokens before scoring (max_doc_tokens=1024). No filtering is applied to the tables.
configsquerieshard negativesteacher scores
hard-negatives · teacher-scores30,000 (all)2,970,047 (2,970,047 dense)3,000,047
python
from datasets import load_dataset
negatives = load_dataset("Hyukkyu/train-text2sql", "hard-negatives", split="train")
scores    = load_dataset("Hyukkyu/train-text2sql", "teacher-scores", split="train")

Jev judgments

judgments holds, for every query of the training sample (the queries with teacher scores), whether TypeSafe's Jev (jev-1.13.0) judged its training positive and its mined candidates relevant: p_yes is Jev's P(yes) for the source's question (e.g. does the passage answer the query?). They locate the false negatives among the mined candidates and the mislabelled positives.

  • —Requests. One request per query (round 0): its training positive, its candidates whose teacher score taken as (cos + 1) / 2 is at least 0.85 × the positive's (at most 12, the highest scores) and 4 random candidates below that, shuffled under neutral ids, one yes/no question per passage. Queries left with fewer than 10 candidates under their source's cutoff got their next hardest unjudged candidates in rounds 1–8 (8 per request), those still under 10 in rounds 9–10 (24 per request). The dataset's own labelled negatives were never sent. Texts were cut to 512 (query) and 512 (passage) tokens of the jina-embeddings-v5 small tokenizer. Jev answers a request's passages in one context, so P(yes) is calibrated to these groups: the thresholds below apply to this table, not to single-pair calls.
  • —Accuracy (an audit of 1237 pairs from the pilot's first requests (100 queries per source; the five long-query sources re-piloted at 512-token queries), labelled blind by an LLM (Claude), at the pilot's fixed thresholds 0.35 and 0.15): a candidate at P(yes) ≥ 0.35 was relevant 67% of the time inside the band (n = 350) and 45% below it (n = 87); one under 0.35 was relevant 6% (band, n = 387) and 1% (below the band, n = 210) of the time. A positive under 0.15 was mislabelled 100% of the time (n = 20) in the sources that keep the check; in dom-casehold, dom-clerc, dom-cornstack-py, dom-finqa10k, dom-gerlayqa, dom-investopedia, dom-lawse, dom-magicoder, dom-medmcqa, dom-pubmedqa, dom-s2orc, dom-tatqa, Jev's flags were right less often (0%–57% in this audit), under the 70% the check needs, so their positives are not checked.
  • —Use (the SPARSE loader, annotation.filter.judge): a candidate at P(yes) ≥ its source's cutoff (below; fitted on 1,521 labelled pairs) is never a negative; a positive under 0.15 is replaced by the candidate Jev scores highest if that is ≥ 0.8, else the query is dropped; a candidate at ≥ 0.9 can become an extra positive. Compare p_yes as a float64 (it is stored as one).
configqueriesrowscandidates per querytop-up rowscandidate cutoffpositive check
judgments30,000399,53012.3215,6240.34< 0.15

Load it

python
from datasets import load_dataset
queries   = load_dataset("Hyukkyu/train-text2sql", "queries", split="train")
corpus    = load_dataset("Hyukkyu/train-text2sql", "corpus", split="train")
qrels     = load_dataset("Hyukkyu/train-text2sql", "qrels", split="train")
negatives = load_dataset("Hyukkyu/train-text2sql", "hard-negatives", split="train")
scores    = load_dataset("Hyukkyu/train-text2sql", "teacher-scores", split="train")
judgments = load_dataset("Hyukkyu/train-text2sql", "judgments", split="train")

License and attribution

The data is redistributed under the source's terms — apache-2.0. All credit belongs to the original authors; see the source repository (https://huggingface.co/datasets/gretelai/synthetictextto_sql). This repository is an independent repackaging.