Team Ai
Datasetpublic

epfl-dlab/JSONSchemaBench

JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the dataset all_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.

sourceHugging Facemitupdated 2y agoView on Hugging Face
12likes3.5kdownloads
README.md430 linesDownload Raw Back to root
1---2pretty_name: J3dataset_info:4- config_name: Github_easy5  features:6  - name: json_schema7    dtype: string8  - name: unique_id9    dtype: string10  splits:11  - name: train12    num_bytes: 120863613    num_examples: 117014  - name: val15    num_bytes: 18268816    num_examples: 19117  - name: test18    num_bytes: 539656.019    num_examples: 57720  download_size: 54061021  dataset_size: 1930980.022- config_name: Github_hard23  features:24  - name: json_schema25    dtype: string26  - name: unique_id27    dtype: string28  splits:29  - name: train30    num_bytes: 1281615231    num_examples: 74632  - name: val33    num_bytes: 160752534    num_examples: 12235  - name: test36    num_bytes: 5754647.48387096737    num_examples: 36838  download_size: 356214639  dataset_size: 20178324.4838709740- config_name: Github_medium41  features:42  - name: json_schema43    dtype: string44  - name: unique_id45    dtype: string46  splits:47  - name: train48    num_bytes: 499083249    num_examples: 118950  - name: val51    num_bytes: 55739052    num_examples: 19453  - name: test54    num_bytes: 2417201.578414839755    num_examples: 58656  download_size: 158033657  dataset_size: 7965423.5784148458- config_name: Github_trivial59  features:60  - name: json_schema61    dtype: string62  - name: unique_id63    dtype: string64  splits:65  - name: train66    num_bytes: 467333.2432432432567    num_examples: 26668  - name: val69    num_bytes: 77303.2432432432470    num_examples: 4471  - name: test72    num_bytes: 235423.5135135135273    num_examples: 13474  download_size: 15804475  dataset_size: 780060.076- config_name: Github_ultra77  features:78  - name: json_schema79    dtype: string80  - name: unique_id81    dtype: string82  splits:83  - name: train84    num_bytes: 7311744.74390243985    num_examples: 9886  - name: val87    num_bytes: 1193754.24390243988    num_examples: 1689  - name: test90    num_bytes: 3730482.01219512291    num_examples: 5092  download_size: 222145593  dataset_size: 12235981.094- config_name: Glaiveai2K95  features:96  - name: json_schema97    dtype: string98  - name: unique_id99    dtype: string100  splits:101  - name: train102    num_bytes: 865943.3989455184103    num_examples: 1026104  - name: val105    num_bytes: 141791.9015817223106    num_examples: 168107  - name: test108    num_bytes: 432971.6994727592109    num_examples: 513110  download_size: 284264111  dataset_size: 1440707.0112- config_name: JsonSchemaStore113  features:114  - name: json_schema115    dtype: string116  - name: unique_id117    dtype: string118  splits:119  - name: train120    num_bytes: 13308367.977642277121    num_examples: 295122  - name: val123    num_bytes: 2210542.4776422763124    num_examples: 49125  - name: test126    num_bytes: 6676740.544715447127    num_examples: 148128  download_size: 4019966129  dataset_size: 22195651.0130- config_name: Kubernetes131  features:132  - name: json_schema133    dtype: string134  - name: unique_id135    dtype: string136  splits:137  - name: train138    num_bytes: 15388503.69924812139    num_examples: 639140  - name: val141    num_bytes: 2528627.3684210526142    num_examples: 105143  - name: test144    num_bytes: 7706292.932330827145    num_examples: 320146  download_size: 6819424147  dataset_size: 25623424.0148- config_name: Snowplow149  features:150  - name: json_schema151    dtype: string152  - name: unique_id153    dtype: string154  splits:155  - name: train156    num_bytes: 969083.2952853598157    num_examples: 242158  - name: val159    num_bytes: 160179.0570719603160    num_examples: 40161  - name: test162    num_bytes: 484541.6476426799163    num_examples: 121164  download_size: 298277165  dataset_size: 1613804.0166- config_name: WashingtonPost167  features:168  - name: json_schema169    dtype: string170  - name: unique_id171    dtype: string172  splits:173  - name: train174    num_bytes: 1604526.016175    num_examples: 74176  - name: val177    num_bytes: 281876.192178    num_examples: 13179  - name: test180    num_bytes: 823945.792181    num_examples: 38182  download_size: 565170183  dataset_size: 2710348.0184- config_name: default185  features:186  - name: json_schema187    dtype: string188  - name: unique_id189    dtype: string190  splits:191  - name: train192    num_bytes: 54520620193    num_examples: 5754194  - name: val195    num_bytes: 15255546196    num_examples: 937197  - name: test198    num_bytes: 27031812.394351464199    num_examples: 2867200  download_size: 20765998201  dataset_size: 96807978.39435147202configs:203- config_name: Github_easy204  data_files:205  - split: train206    path: Github_easy/train-*207  - split: val208    path: Github_easy/val-*209  - split: test210    path: Github_easy/test-*211- config_name: Github_hard212  data_files:213  - split: train214    path: Github_hard/train-*215  - split: val216    path: Github_hard/val-*217  - split: test218    path: Github_hard/test-*219- config_name: Github_medium220  data_files:221  - split: train222    path: Github_medium/train-*223  - split: val224    path: Github_medium/val-*225  - split: test226    path: Github_medium/test-*227- config_name: Github_trivial228  data_files:229  - split: train230    path: Github_trivial/train-*231  - split: val232    path: Github_trivial/val-*233  - split: test234    path: Github_trivial/test-*235- config_name: Github_ultra236  data_files:237  - split: train238    path: Github_ultra/train-*239  - split: val240    path: Github_ultra/val-*241  - split: test242    path: Github_ultra/test-*243- config_name: Glaiveai2K244  data_files:245  - split: train246    path: Glaiveai2K/train-*247  - split: val248    path: Glaiveai2K/val-*249  - split: test250    path: Glaiveai2K/test-*251- config_name: JsonSchemaStore252  data_files:253  - split: train254    path: JsonSchemaStore/train-*255  - split: val256    path: JsonSchemaStore/val-*257  - split: test258    path: JsonSchemaStore/test-*259- config_name: Kubernetes260  data_files:261  - split: train262    path: Kubernetes/train-*263  - split: val264    path: Kubernetes/val-*265  - split: test266    path: Kubernetes/test-*267- config_name: Snowplow268  data_files:269  - split: train270    path: Snowplow/train-*271  - split: val272    path: Snowplow/val-*273  - split: test274    path: Snowplow/test-*275- config_name: WashingtonPost276  data_files:277  - split: train278    path: WashingtonPost/train-*279  - split: val280    path: WashingtonPost/val-*281  - split: test282    path: WashingtonPost/test-*283- config_name: default284  data_files:285  - split: train286    path: data/train-*287  - split: val288    path: data/val-*289  - split: test290    path: data/test-*291license: mit292task_categories:293- text-generation294---295 296# JSONSchemaBench297 298[![Paper](https://img.shields.io/badge/Paper-arXiv-blue)](https://arxiv.org/abs/2501.10868)299[![GitHub](https://img.shields.io/badge/Code-GitHub-blue)](https://github.com/guidance-ai/jsonschemabench)300 301JSONSchemaBench is a benchmark of **real-world JSON schemas** designed to evaluate **structured output generation** for Large Language Models (LLMs). It contains approximately **10,000 JSON schemas**, capturing diverse constraints and complexities.302 303 304```python305import datasets306from datasets import load_dataset307 308def main():309    # Inspect the available subsets of the dataset310    all_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")311    print("Available subsets:", all_subsets)312    # Example output: ['Github_easy', 'Github_hard', 'Github_medium', 'Github_trivial', 'Github_ultra', 'Glaiveai2K', 'JsonSchemaStore', 'Kubernetes', 'Snowplow', 'WashingtonPost', 'default']313 314    # Access a specific subset of the dataset315    subset_name = "Github_easy"316    github_easy = load_dataset("epfl-dlab/JSONSchemaBench", subset_name)317    print(f"Loaded subset '{subset_name}':", github_easy)318 319    # Load the entire dataset as a whole320    entire_dataset = load_dataset("epfl-dlab/JSONSchemaBench", "default")321    print("Loaded entire dataset:", entire_dataset)322 323if __name__ == "__main__":324    main()325```326 327## Update (March 31st, 2025)328 329To improve inference efficiency and streamline data collation, we’ve decided to drop a small number of exceptionally long samples from the dataset.330 331We’re using the `meta-llama/Llama-3.2-1B-instruct` tokenizer, and the filtering criteria are as follows:332- Github_easy: Samples longer than 1024 tokens — 5 out of 582 removed333- Github_medium: Samples longer than 2048 tokens — 7 out of 593 removed334- Github_hard: Samples longer than 8192 tokens — 4 out of 372 removed335- Other subsets are not touched336 337Since the number of discarded samples is minimal, this change is expected to have at most a 1% impact on results.338 339 340## ⚠️ Important Update (March 10th, 2025)341 342We have restructured the dataset to include train/val/test splits. If you downloaded the dataset before this date, you might encounter errors like `KeyError: 'Github_easy'`.343 344To fix this issue, please follow one of the options below:345 3461. Update How Subsets Are Accessed:347If you previously used:348 349```python350from datasets import load_dataset, concatenate_datasets, DatasetDict, Dataset351 352subset: DatasetDict = load_dataset("epfl-dlab/JSONSchemaBench")353subset["Github_easy"]354```355You can update it to:356 357```python358from datasets import load_dataset, concatenate_datasets, DatasetDict, Dataset359 360subset: DatasetDict = load_dataset("epfl-dlab/JSONSchemaBench", name="Github_easy")361subset: Dataset = concatenate_datasets([subset["train"], subset["val"], subset["test"]])362```363 3642. Load the Dataset in the Old Structure:365If you need the previous structure, you can use a specific revision:366 367```python368dataset = load_dataset("epfl-dlab/JSONSchemaBench", revision="e2ee5fdba65657c60d3a24b321172eb7141f8d73")369```370 371We apologize for the inconvenience and appreciate your understanding! 😊372 373## 📌 Dataset Overview374- **Purpose:** Evaluate the **efficiency** and **coverage** of structured output generation.375- **Sources:** GitHub, Kubernetes, API specifications, curated collections.376- **Schemas:** Categorized based on complexity and domain.377 378### 📊 Dataset Breakdown379| Dataset         | Category            | Count |380| --------------- | ------------------- | ----- |381| GlaiveAI-2K     | Function Call       | 1707  |382| Github-Trivial  | Misc                | 444   |383| Github-Easy     | Misc                | 1943  |384| Snowplow        | Operational API     | 403   |385| Github-Medium   | Misc                | 1976  |386| Kubernetes      | Kubernetes API      | 1064  |387| Washington Post | Resource Access API | 125   |388| Github-Hard     | Misc                | 1240  |389| JSONSchemaStore | Misc                | 492   |390| Github-Ultra    | Misc                | 164   |391| **Total**       |                     | 9558  |392 393## 📥 Loading the Dataset394 395```python396from datasets import load_dataset397 398dataset = load_dataset("epfl-dlab/JSONSchemaBench")399print(dataset)400```401 402## 🔍 Data Structure403Each dataset split contains:404- `"json_schema"`: The schema definition.405- `"unique_id"`: A unique identifier for the schema.406 407 408🚀 **For more details, check out the [paper](https://arxiv.org/abs/2501.10868).**409 410## 📚 Citation411```bibtex412@misc{geng2025jsonschemabench,413      title={Generating Structured Outputs from Language Models: Benchmark and Studies},414      author={Saibo Geng et al.},415      year={2025},416      eprint={2501.10868},417      archivePrefix={arXiv},418      primaryClass={cs.CL},419      url={https://arxiv.org/abs/2501.10868}420}421```422 423 424## License425 426This dataset is provided under the [MIT License](https://opensource.org/licenses/MIT). Please ensure that you comply with the license terms when using or distributing this dataset.427 428## Acknowledgements429 430We would like to thank the contributors and maintainers of the JSON schema projects and the open-source community for their invaluable work and support.