AriaAICompany/phish-messages
Phish synthetic messages 256 synthetic Persian and English messages for the Phish review demo. Seed 3. Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy. Files data/messages.jsonl data/splits.json data/evaluation.json data/sample_preview.json data/eml/*.eml data/protocol.md Splits Split is by campaign… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/phish-messages.
Phish synthetic messages
256 synthetic Persian and English messages for the Phish review demo. Seed 3.
Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy.
Files
- data/messages.jsonl
- data/splits.json
- data/evaluation.json
- data/sample_preview.json
- data/eml/*.eml
- data/protocol.md
Splits
Split is by campaign, sender group, and template, not random rows of the same lure.
- development: 10 campaigns × 16 variants = 160
- test: 6 campaigns × 16 variants = 96
- interactive: six
*-00fixtures used in the Space
Phishing, hard-negative invoices/login notices, and insufficient-evidence stubs are all present. Hard negatives keep aligned From/Reply-To/href domains.
Labels
gold_binary: 1 = phishing campaign, 0 = hard negative or insufficientgold_label:suspiciousorinsufficient_evidence- The demo never emits
safe/benign
License
CC-BY-4.0. Keep the synthetic-data label. Do not treat hosts under .example as live infrastructure.
