Team Ai
Datasetpublic

316usman/email-security

EMAIL_SECURITY A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.

sourceHugging Faceunknownupdated 8d agoView on Hugging Face
0likes102downloads
Dataset Card

EMAIL_SECURITY

A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.

Format

Standard preference / DPO schema — each row:

columnmeaning
promptthe request (originally prompt)
chosenthe human-preferred response
rejecteda worse response to the same prompt
sourcethe dataset/URL the row was harvested from

Splits

80/10/10 train / validation / test (seeded shuffle): train:1296 / validation:162 / test:163

Stats

  • —Rows: 1621
  • —Distinct sources: 1

Sources

  • —synthetic:deepseek/deepseek-v4-flash-0731

Provenance

`prompt` and `chosen` are real — harvested from the human-labelled sources listed above (accepted / upvoted / high-rated responses).

`rejected` is partly synthetic. Most sources retain only the accepted answer, so where no genuine low-scored counterpart existed the negative was generated by defect injection: an LLM rewrites that row's own chosen with exactly one flaw introduced (omission, factual error, wrong remedy, unsupported claim, or misprioritisation) at matched length. Rows carrying a defect column are the generated ones; rows without it kept a negative that was present in the source.

Rows passed a per-row check (negative is unique, not a stub, not spliced from chosen) and a whole-corpus gate for repeated negatives, length shortcuts, train/eval leakage and bag-of-words separability.