epfl-dlab/JSONSchemaBench
JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the dataset all_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.
123.5k
1---2pretty_name: J3dataset_info:4- config_name: Github_easy5 features:6 - name: json_schema7 dtype: string8 - name: unique_id9 dtype: string10 splits:11 - name: train12 num_bytes: 120863613 num_examples: 117014 - name: val15 num_bytes: 18268816 num_examples: 19117 - name: test18 num_bytes: 539656.019 num_examples: 57720 download_size: 54061021 dataset_size: 1930980.022- config_name: Github_hard23 features:24 - name: json_schema25 dtype: string26 - name: unique_id27 dtype: string28 splits:29 - name: train30 num_bytes: 1281615231 num_examples: 74632 - name: val33 num_bytes: 160752534 num_examples: 12235 - name: test36 num_bytes: 5754647.48387096737 num_examples: 36838 download_size: 356214639 dataset_size: 20178324.4838709740- config_name: Github_medium41 features:42 - name: json_schema43 dtype: string44 - name: unique_id45 dtype: string46 splits:47 - name: train48 num_bytes: 499083249 num_examples: 118950 - name: val51 num_bytes: 55739052 num_examples: 19453 - name: test54 num_bytes: 2417201.578414839755 num_examples: 58656 download_size: 158033657 dataset_size: 7965423.5784148458- config_name: Github_trivial59 features:60 - name: json_schema61 dtype: string62 - name: unique_id63 dtype: string64 splits:65 - name: train66 num_bytes: 467333.2432432432567 num_examples: 26668 - name: val69 num_bytes: 77303.2432432432470 num_examples: 4471 - name: test72 num_bytes: 235423.5135135135273 num_examples: 13474 download_size: 15804475 dataset_size: 780060.076- config_name: Github_ultra77 features:78 - name: json_schema79 dtype: string80 - name: unique_id81 dtype: string82 splits:83 - name: train84 num_bytes: 7311744.74390243985 num_examples: 9886 - name: val87 num_bytes: 1193754.24390243988 num_examples: 1689 - name: test90 num_bytes: 3730482.01219512291 num_examples: 5092 download_size: 222145593 dataset_size: 12235981.094- config_name: Glaiveai2K95 features:96 - name: json_schema97 dtype: string98 - name: unique_id99 dtype: string100 splits:101 - name: train102 num_bytes: 865943.3989455184103 num_examples: 1026104 - name: val105 num_bytes: 141791.9015817223106 num_examples: 168107 - name: test108 num_bytes: 432971.6994727592109 num_examples: 513110 download_size: 284264111 dataset_size: 1440707.0112- config_name: JsonSchemaStore113 features:114 - name: json_schema115 dtype: string116 - name: unique_id117 dtype: string118 splits:119 - name: train120 num_bytes: 13308367.977642277121 num_examples: 295122 - name: val123 num_bytes: 2210542.4776422763124 num_examples: 49125 - name: test126 num_bytes: 6676740.544715447127 num_examples: 148128 download_size: 4019966129 dataset_size: 22195651.0130- config_name: Kubernetes131 features:132 - name: json_schema133 dtype: string134 - name: unique_id135 dtype: string136 splits:137 - name: train138 num_bytes: 15388503.69924812139 num_examples: 639140 - name: val141 num_bytes: 2528627.3684210526142 num_examples: 105143 - name: test144 num_bytes: 7706292.932330827145 num_examples: 320146 download_size: 6819424147 dataset_size: 25623424.0148- config_name: Snowplow149 features:150 - name: json_schema151 dtype: string152 - name: unique_id153 dtype: string154 splits:155 - name: train156 num_bytes: 969083.2952853598157 num_examples: 242158 - name: val159 num_bytes: 160179.0570719603160 num_examples: 40161 - name: test162 num_bytes: 484541.6476426799163 num_examples: 121164 download_size: 298277165 dataset_size: 1613804.0166- config_name: WashingtonPost167 features:168 - name: json_schema169 dtype: string170 - name: unique_id171 dtype: string172 splits:173 - name: train174 num_bytes: 1604526.016175 num_examples: 74176 - name: val177 num_bytes: 281876.192178 num_examples: 13179 - name: test180 num_bytes: 823945.792181 num_examples: 38182 download_size: 565170183 dataset_size: 2710348.0184- config_name: default185 features:186 - name: json_schema187 dtype: string188 - name: unique_id189 dtype: string190 splits:191 - name: train192 num_bytes: 54520620193 num_examples: 5754194 - name: val195 num_bytes: 15255546196 num_examples: 937197 - name: test198 num_bytes: 27031812.394351464199 num_examples: 2867200 download_size: 20765998201 dataset_size: 96807978.39435147202configs:203- config_name: Github_easy204 data_files:205 - split: train206 path: Github_easy/train-*207 - split: val208 path: Github_easy/val-*209 - split: test210 path: Github_easy/test-*211- config_name: Github_hard212 data_files:213 - split: train214 path: Github_hard/train-*215 - split: val216 path: Github_hard/val-*217 - split: test218 path: Github_hard/test-*219- config_name: Github_medium220 data_files:221 - split: train222 path: Github_medium/train-*223 - split: val224 path: Github_medium/val-*225 - split: test226 path: Github_medium/test-*227- config_name: Github_trivial228 data_files:229 - split: train230 path: Github_trivial/train-*231 - split: val232 path: Github_trivial/val-*233 - split: test234 path: Github_trivial/test-*235- config_name: Github_ultra236 data_files:237 - split: train238 path: Github_ultra/train-*239 - split: val240 path: Github_ultra/val-*241 - split: test242 path: Github_ultra/test-*243- config_name: Glaiveai2K244 data_files:245 - split: train246 path: Glaiveai2K/train-*247 - split: val248 path: Glaiveai2K/val-*249 - split: test250 path: Glaiveai2K/test-*251- config_name: JsonSchemaStore252 data_files:253 - split: train254 path: JsonSchemaStore/train-*255 - split: val256 path: JsonSchemaStore/val-*257 - split: test258 path: JsonSchemaStore/test-*259- config_name: Kubernetes260 data_files:261 - split: train262 path: Kubernetes/train-*263 - split: val264 path: Kubernetes/val-*265 - split: test266 path: Kubernetes/test-*267- config_name: Snowplow268 data_files:269 - split: train270 path: Snowplow/train-*271 - split: val272 path: Snowplow/val-*273 - split: test274 path: Snowplow/test-*275- config_name: WashingtonPost276 data_files:277 - split: train278 path: WashingtonPost/train-*279 - split: val280 path: WashingtonPost/val-*281 - split: test282 path: WashingtonPost/test-*283- config_name: default284 data_files:285 - split: train286 path: data/train-*287 - split: val288 path: data/val-*289 - split: test290 path: data/test-*291license: mit292task_categories:293- text-generation294---295 296# JSONSchemaBench297 298[](https://arxiv.org/abs/2501.10868)299[](https://github.com/guidance-ai/jsonschemabench)300 301JSONSchemaBench is a benchmark of **real-world JSON schemas** designed to evaluate **structured output generation** for Large Language Models (LLMs). It contains approximately **10,000 JSON schemas**, capturing diverse constraints and complexities.302 303 304```python305import datasets306from datasets import load_dataset307 308def main():309 # Inspect the available subsets of the dataset310 all_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")311 print("Available subsets:", all_subsets)312 # Example output: ['Github_easy', 'Github_hard', 'Github_medium', 'Github_trivial', 'Github_ultra', 'Glaiveai2K', 'JsonSchemaStore', 'Kubernetes', 'Snowplow', 'WashingtonPost', 'default']313 314 # Access a specific subset of the dataset315 subset_name = "Github_easy"316 github_easy = load_dataset("epfl-dlab/JSONSchemaBench", subset_name)317 print(f"Loaded subset '{subset_name}':", github_easy)318 319 # Load the entire dataset as a whole320 entire_dataset = load_dataset("epfl-dlab/JSONSchemaBench", "default")321 print("Loaded entire dataset:", entire_dataset)322 323if __name__ == "__main__":324 main()325```326 327## Update (March 31st, 2025)328 329To improve inference efficiency and streamline data collation, we’ve decided to drop a small number of exceptionally long samples from the dataset.330 331We’re using the `meta-llama/Llama-3.2-1B-instruct` tokenizer, and the filtering criteria are as follows:332- Github_easy: Samples longer than 1024 tokens — 5 out of 582 removed333- Github_medium: Samples longer than 2048 tokens — 7 out of 593 removed334- Github_hard: Samples longer than 8192 tokens — 4 out of 372 removed335- Other subsets are not touched336 337Since the number of discarded samples is minimal, this change is expected to have at most a 1% impact on results.338 339 340## ⚠️ Important Update (March 10th, 2025)341 342We have restructured the dataset to include train/val/test splits. If you downloaded the dataset before this date, you might encounter errors like `KeyError: 'Github_easy'`.343 344To fix this issue, please follow one of the options below:345 3461. Update How Subsets Are Accessed:347If you previously used:348 349```python350from datasets import load_dataset, concatenate_datasets, DatasetDict, Dataset351 352subset: DatasetDict = load_dataset("epfl-dlab/JSONSchemaBench")353subset["Github_easy"]354```355You can update it to:356 357```python358from datasets import load_dataset, concatenate_datasets, DatasetDict, Dataset359 360subset: DatasetDict = load_dataset("epfl-dlab/JSONSchemaBench", name="Github_easy")361subset: Dataset = concatenate_datasets([subset["train"], subset["val"], subset["test"]])362```363 3642. Load the Dataset in the Old Structure:365If you need the previous structure, you can use a specific revision:366 367```python368dataset = load_dataset("epfl-dlab/JSONSchemaBench", revision="e2ee5fdba65657c60d3a24b321172eb7141f8d73")369```370 371We apologize for the inconvenience and appreciate your understanding! 😊372 373## 📌 Dataset Overview374- **Purpose:** Evaluate the **efficiency** and **coverage** of structured output generation.375- **Sources:** GitHub, Kubernetes, API specifications, curated collections.376- **Schemas:** Categorized based on complexity and domain.377 378### 📊 Dataset Breakdown379| Dataset | Category | Count |380| --------------- | ------------------- | ----- |381| GlaiveAI-2K | Function Call | 1707 |382| Github-Trivial | Misc | 444 |383| Github-Easy | Misc | 1943 |384| Snowplow | Operational API | 403 |385| Github-Medium | Misc | 1976 |386| Kubernetes | Kubernetes API | 1064 |387| Washington Post | Resource Access API | 125 |388| Github-Hard | Misc | 1240 |389| JSONSchemaStore | Misc | 492 |390| Github-Ultra | Misc | 164 |391| **Total** | | 9558 |392 393## 📥 Loading the Dataset394 395```python396from datasets import load_dataset397 398dataset = load_dataset("epfl-dlab/JSONSchemaBench")399print(dataset)400```401 402## 🔍 Data Structure403Each dataset split contains:404- `"json_schema"`: The schema definition.405- `"unique_id"`: A unique identifier for the schema.406 407 408🚀 **For more details, check out the [paper](https://arxiv.org/abs/2501.10868).**409 410## 📚 Citation411```bibtex412@misc{geng2025jsonschemabench,413 title={Generating Structured Outputs from Language Models: Benchmark and Studies},414 author={Saibo Geng et al.},415 year={2025},416 eprint={2501.10868},417 archivePrefix={arXiv},418 primaryClass={cs.CL},419 url={https://arxiv.org/abs/2501.10868}420}421```422 423 424## License425 426This dataset is provided under the [MIT License](https://opensource.org/licenses/MIT). Please ensure that you comply with the license terms when using or distributing this dataset.427 428## Acknowledgements429 430We would like to thank the contributors and maintainers of the JSON schema projects and the open-source community for their invaluable work and support.