Team Ai
Datasetpublic

robworks-software/database-query-logs-synthetic

Database Query Logs (synthetic) 3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata. These queries are synthetic The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.

sourceHugging Facecc0-1.0updated 3mo agoView on Hugging Face
0likes51downloads
Dataset Card

Database Query Logs (synthetic)

3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata.

These queries are synthetic

The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema names.

The generation shows through:

  • —3,195 training rows contain 2,243 distinct queries (70.2%); one Elasticsearch query body appears 35 times.
  • —The same query recurs with only the schema name swapped - DELETE FROM telecom_logs and DELETE FROM construction_logs are otherwise byte-identical.
  • —Some generated SQL is not valid. Schema names containing hyphens appear unquoted (DELETE FROM E-COMMERCE.AUDIT_LOGS), which will not parse.

A previous version of this card described the contents as "real-world and synthetic". There is no verified production-captured subset; treat the whole dataset as synthetic.

Loading

python
from datasets import load_dataset

ds = load_dataset("robworks-software/database-query-logs-synthetic")

Splits

SplitRows
train3,195
validation396
test404
total3,995

Appropriate use

  • —Parser and tokenizer testing across dialects (bearing in mind some queries are invalid).
  • —Query-type and complexity classification.
  • —Teaching examples of dialect syntax differences.

Inappropriate use

  • —Performance modeling or optimizer research. execution_time_ms, rows_examined, and rows_returned are generated values. They were not measured on any real system and have no relationship to what these queries would actually cost. Do not train a cost model on them.
  • —Workload characterization. The distribution reflects the generator's template mix, not any real application's query pattern.

Limitations

  • —Synthetic throughout - see above.
  • —Some queries are syntactically invalid.
  • —70.2% distinct queries, with heavy near-duplication beyond exact repeats.
  • —Timing and row-count fields are fabricated.
  • —Schema names are industry labels (TELECOMMUNICATIONS, EDUCATION, LOGISTICS) applied cosmetically; they do not reflect real domain schemas.

License

CC0-1.0. Entirely generated content, dedicated to the public domain.

Citation

bibtex
@dataset{database_query_logs_synthetic,
  title  = {Database Query Logs (synthetic)},
  author = {Robworks Software},
  year   = {2025},
  publisher = {Hugging Face},
  note   = {Programmatically generated queries; timing metadata is fabricated},
  url    = {https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic}
}