unalignment
toxic-dpo-v0.2
Toxic-DPO
This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples.
Many of the examples still contain some amount of warnings/disclaimers, so it's still somewhat editorialized.
Usage restriction
To use this data, you must acknowledge/agree to the following:
data contained within is "toxic"/"harmful", and contains profanity and other types… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/toxic-dpo-v0.2.toxic-dpo-v0.1
Toxic-DPO
This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples.
Most of the examples still contain some amount of warnings/disclaimers, so it's still somewhat editorialized.
Usage restriction
To use this data, you must acknowledge/agree to the following:
data contained within is "toxic"/"harmful", and contains profanity and other types… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/toxic-dpo-v0.1.comedy-snippets-v0.1A very small sampling of snippets of comedy routines by George Carlin and Tom Segura.
meissa-unalignments
Meissa Unalignments
During the Meissa-Qwen2.5 training process, I noticed that Qwen's censorship was somehow bound to its chat template. Thus, I created this dataset with one system prompt, hoping to make it more effective in uncensoring models.
The dataset consists of:
V3N0M/Jenna-50K-Alpaca-Uncensored
jondurbin/airoboros-3.2, category = unalignment
Orion-zhen/dpo-toxic-zh, prompt and chosen
NobodyExistsOnTheInternet/ToxicQAFinal
spicy-3.1Airoboros 3.1 dataset with the spicy/decensorship data re-added.
unalignment-toxic-dpo-v0.1
