Team Ai
Datasetpublic

StanzaAPI/lei-test-data

LEI Test Data Twenty-character LEIs with a correctly computed ISO 7064 MOD-97-10 check-digit pair. Synthetic entities; not registered organisations. Free to use under CC0-1.0 — public domain dedication, no attribution required. Files lei-small.json / lei-small.csv — 100 rows (documents: 25) lei-medium.json / lei-medium.csv — 2,000 rows (documents: 250) lei-large.json / lei-large.csv — 20,000 rows (documents: 2,500) Deterministic: the same size always produces… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/lei-test-data.

sourceHugging Facecc0-1.0updated 6d agoView on Hugging Face
0likes318downloads
Dataset Card

LEI Test Data

Twenty-character LEIs with a correctly computed ISO 7064 MOD-97-10 check-digit pair. Synthetic entities; not registered organisations.

Free to use under CC0-1.0 — public domain dedication, no attribution required.

Files

  • —lei-small.json / lei-small.csv — 100 rows (documents: 25)
  • —lei-medium.json / lei-medium.csv — 2,000 rows (documents: 250)
  • —lei-large.json / lei-large.csv — 20,000 rows (documents: 2,500)

Deterministic: the same size always produces the same rows.

Source

Generated by StanzaAPI from <https://stanzaapi.com/datasets/lei>.

Citation

StanzaAPI. (2026). LEI Test Data. CC0-1.0.
DOI: 10.5281/zenodo.23076180
https://stanzaapi.com/datasets/lei

Security: adversarial corpus

This dataset ships a catalogued adversarial corpus of 12 hostile inputs (injection, XXE, resource exhaustion, encoding attacks, validation bypass, structural confusion, and prompt injection for AI agents) with the safe expected behaviour. See `adversarial/` or the hub.