316usman/email-security
EMAIL_SECURITY A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.
EMAIL_SECURITY
A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
Splits
80/10/10 train / validation / test (seeded shuffle): train:1296 / validation:162 / test:163
Stats
- Rows: 1621
- Distinct sources: 1
Sources
synthetic:deepseek/deepseek-v4-flash-0731
Provenance
`prompt` and `chosen` are real — harvested from the human-labelled sources listed above (accepted / upvoted / high-rated responses).
`rejected` is partly synthetic. Most sources retain only the accepted answer, so where no genuine low-scored counterpart existed the negative was generated by defect injection: an LLM rewrites that row's own chosen with exactly one flaw introduced (omission, factual error, wrong remedy, unsupported claim, or misprioritisation) at matched length. Rows carrying a defect column are the generated ones; rows without it kept a negative that was present in the source.
Rows passed a per-row check (negative is unique, not a stub, not spliced from chosen) and a whole-corpus gate for repeated negatives, length shortcuts, train/eval leakage and bag-of-words separability.
