Team Ai
Datasetpublic

Graunt/json-repair-eval-sample

JSON repair eval (sample) 30 cases of broken JSON. Each one has the text exactly as a parser would receive it, the repair we expect, the breakage category, the rule applied and the reason. It's a sample of a 300-case set for testing the repair step that sits behind an LLM's structured output or a stream that got cut off. There are ten categories: truncation, trailing commas, single quotes, unescaped control characters, NaN and Infinity, comments, concatenated objects, unquoted… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/json-repair-eval-sample.

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes45downloads
Dataset Card

JSON repair eval (sample)

30 cases of broken JSON. Each one has the text exactly as a parser would receive it, the repair we expect, the breakage category, the rule applied and the reason. It's a sample of a 300-case set for testing the repair step that sits behind an LLM's structured output or a stream that got cut off.

There are ten categories: truncation, trailing commas, single quotes, unescaped control characters, NaN and Infinity, comments, concatenated objects, unquoted keys, mismatched brackets, and mixed faults. This sample has three cases from each.

Three of the 30 have no correct repair (unrecoverable: true). Take {"invoice_no":"INV-77","total_cents":129: the total could be 129, 1295 or 12950, and the bytes don't say. A repair layer should decline those. Any confident fix counts as a failure.

Scoring

Parse your output and expected_repaired, then compare the values. Don't compare strings, because whitespace and key order mean nothing in JSON. On an unrecoverable case, declining is the only pass.

Fields

fieldmeaning
case_idstable id
broken_inputthe malformed text
expected_repairedthe expected JSON text, or null when there is no correct repair
repair_categoryone of the ten categories, defined in repair_categories.json
applied_policythe rule that produced the repair
rationalewhy that repair, or why none
scoring_noteshow to score the row
unrecoverabletrue when no repair is correct

The full set

All 300 cases (270 with a repair, 30 without) are on Graunt, sold by Lapidary: JSON Repair Eval Set.

Origin

Every case was written for this set from synthetic seeds by a deterministic build program. No language model generated any case, and no value refers to a real person or company. This sample is licensed CC BY 4.0.