Team Ai
Datasetpublic

StanzaAPI/x12-test-data

ANSI X12 Test Data Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files x12-small.json / x12-small.csv — 100 rows (documents: 25) x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250) x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/x12-test-data.

sourceHugging Facecc0-1.0updated 7d agoView on Hugging Face
0likes665downloads
Dataset Card

ANSI X12 Test Data

Structurally complete X12 837P claim documents (ISA, GS, ST, BHT, NM1, CLM, SE, GE, IEA) with varying control numbers. Synthetic claims; not real patients or payers.

Free to use under CC0-1.0 — public domain dedication, no attribution required.

Files

  • —x12-small.json / x12-small.csv — 100 rows (documents: 25)
  • —x12-medium.json / x12-medium.csv — 2,000 rows (documents: 250)
  • —x12-large.json / x12-large.csv — 20,000 rows (documents: 2,500)

Deterministic: the same size always produces the same rows.

Source

Generated by StanzaAPI from <https://stanzaapi.com/datasets/x12>.

Citation

StanzaAPI. (2026). ANSI X12 Test Data. CC0-1.0.
DOI: 10.5281/zenodo.23076202
https://stanzaapi.com/datasets/x12

Security: adversarial corpus

This dataset ships a catalogued adversarial corpus of 12 hostile inputs (injection, XXE, resource exhaustion, encoding attacks, validation bypass, structural confusion, and prompt injection for AI agents) with the safe expected behaviour. See `adversarial/` or the hub.