text-2-sql
synthetic-text2sqlEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment:
import mteb
import logging
from sentence_transformers import SentenceTransformer
from mteb import MTEB
logger = logging.getLogger(__name__)
model_name = 'intfloat/e5-base-v2'
model = SentenceTransformer(model_name)
tasks = mteb.get_tasks(
tasks=[
"AppsRetrieval",
"CodeFeedbackMT",
"CodeFeedbackST",
"CodeTransOceanContest",
"CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql.SQaLe-2-text-to-SQL-SchemasSQaLe: schemas and databases
Project page ·
Questions and SQL ·
Trained models ·
Python library ·
Citation
This dataset holds the 9,259 populated databases of SQaLe, a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. Each row is one database: its DDL, extended from a real schema in SchemaPile, and the generated rows of its tables. The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas.text2sql-loop-engineering-data
text2sql-loop-engineering course data
Data package for the course repository https://github.com/mushan-shine/text2sql-loop-engineering .
Load it into your own Databricks workspace from a notebook with scripts/load_course_data.py.
Contents: the BEAVER dw data warehouse (97 tables) and the benchmark tables prepared in step 1
(questions, table metadata, gold results validated against BEAVER's official MySQL engine), plus the
dev set and step-1 result files.
Source and… See the full description on the dataset page: https://huggingface.co/datasets/MuShan795/text2sql-loop-engineering-data.text2sql-eval-results
Text2SQL Evaluation Toolkit — Pre-computed Results
Pre-computed inference and evaluation results produced by the
IBM/text2sql-eval-toolkit
across six text-to-SQL benchmarks.
These artefacts power the toolkit's evaluation dashboard and analysis scripts
without requiring you to re-run multi-hour inference pipelines.
Quick start
Install the toolkit and fetch all results (~3.8 GB):
pip install text2sql-eval-toolkit
text2sql-eval-toolkit results fetch
Fetch a single… See the full description on the dataset page: https://huggingface.co/datasets/text2sql-eval-toolkit/text2sql-eval-results.SQaLe-2-text-to-SQL-QueriesSQaLe: questions and SQL
Project page ·
Schemas and databases ·
Trained models ·
Python library ·
Citation
SQaLe is a large semi-synthetic text-to-SQL dataset grounded in real-world database schemas, introduced in the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It pairs 1,408,056 natural-language questions with 176,761 distinct SQL queries over 9,259 populated SQLite databases. The schemas come from SchemaPile, a collection of database… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries.RBAC-Text2SQL-Benchmark
RBAC-Text2SQL Benchmark
Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that
respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural
language question with a role policy; the model must either produce a correct SQL query
that touches only authorized resources, or refuse with Sorry, I cannot answer.
Code, evaluation harness, and reproduction instructions:
https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.
