Team Ai
Datasetpublic

llmsql-bench/llmsql-benchmark

LLMSQL Benchmark ⚠️ A newer version of this dataset is available:πŸ‘‰ https://huggingface.co/datasets/llmsql-bench/llmsql-2.0 This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark. Arxiv Article: https://arxiv.org/abs/2510.02350 Files tables.jsonl β€” Database table metadata questions.jsonl β€” All available questions train_questions.jsonl, val_questions.jsonl, test_questions.jsonl β€” Data… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-benchmark.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
2likes217downloads
Dataset Card

LLMSQL Benchmark

⚠️ A newer version of this dataset is available: πŸ‘‰ https://huggingface.co/datasets/llmsql-bench/llmsql-2.0

This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark.

Arxiv Article: https://arxiv.org/abs/2510.02350

Files

  • β€”tables.jsonl β€” Database table metadata
  • β€”questions.jsonl β€” All available questions
  • β€”train_questions.jsonl, val_questions.jsonl, test_questions.jsonl β€” Data splits for finetuning, see https://github.com/LLMSQL/llmsql-benchmark
  • β€”sqlite_tables.db β€” sqlite db with tables from tables.jsonl, created with the help of create_db_sql.
  • β€”create_db.sql β€” SQL script that creates the database sqlite_tables.db.

test_output.jsonl is not included in the dataset.

Citation

If you use this benchmark, please cite:

@inproceedings{llmsql_bench,
  title={LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQLels},
  author={Pihulski, Dzmitry and  Charchut, Karol and Novogrodskaia, Viktoria and Koco{'n}, Jan},
  booktitle={2025 IEEE International Conference on Data Mining Workshops (ICDMW)},
  year={2025},
  organization={IEEE}
}