heretic-org/Multilingual-Harmless-Harmful
Multilingual Harmless and Harmful Prompts What is this? This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 9 languages (including original English data). This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Multilingual-Harmless-Harmful.
Multilingual Harmless and Harmful Prompts
What is this?
This dataset contains the Translations of the (1) `heretic-org/Semantic-Harmless` dataset and the (2) `heretic-org/Semantic-Harmful` dataset into 9 languages (including original English data).
This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can use this dataset for any research related to ["Abliteration in a multilingual model. GitHub #65"](https://github.com/p-e-w/heretic/issues/65) or "How refusal behavior differs across different languages" or any similar topics.
[!NOTE] I speak Hindi, Marathi, and somewhat Italian, so I actually checked the translations for a few of them. But I'm not going to pretend that I'm an expert linguist who can guarantee everything is "perfect". So far, it Looks good to me. The rest of the languages like Bengali, French, Russian, Tamil... I don't really understand any of them, so I can't say how good those translations are. If you find any wrong translations, PRs to fix them are much welcome.
Usage with Heretic
Each language in this dataset has its own config/subset, where each of these configs have two splits: harmless and harmful.
Example using CLI
heretic --model Qwen/Qwen3-VL-8B-Instruct \
--good-prompts.dataset "heretic-org/Multilingual-Harmless-Harmful" \
--good-prompts.config "hindi" \
--good-prompts.split "harmless[:400]" \
--good-prompts.column "text" \
--bad-prompts.dataset "heretic-org/Multilingual-Harmless-Harmful" \
--bad-prompts.config "hindi" \
--bad-prompts.split "harmful[:400]" \
--bad-prompts.column "text"To support specifying a dataset's specific config/subset, it was implemented in feat: Support specifying a dataset's specific config/subset #445 for p-e-w/heretic.
Example using config.toml
# Using first 5 prompts for abliteration.
[good_prompts]
dataset = "heretic-org/Multilingual-Harmless-Harmful"
config = "hindi"
split = "harmless[:5]"
column = "text"
[bad_prompts]
dataset = "heretic-org/Multilingual-Harmless-Harmful"
config = "hindi"
split = "harmful[:5]"
column = "text"
# Using first 6-10 prompts for evaluation.
[scorer.KLDivergence.prompts]
dataset = "heretic-org/Multilingual-Harmless-Harmful"
config = "hindi"
split = "harmless[5:10]"
column = "text"
[scorer.KeywordRate.prompts]
dataset = "heretic-org/Multilingual-Harmless-Harmful"
config = "hindi"
split = "harmful[5:10]"
column = "text"Languages
Extra
For convenience, this dataset is also available in PLAINTEXT (`.txt`), CSV (`.csv`), and JSON (`.json`) formats inside the `content/` folder.
1. Structure:
content/
└── hindi/
├── text/
│ ├── hindi_harmful.txt
│ └── hindi_harmless.txt
├── json/
│ └── hindi.json
└── csv/
└── hindi.csv2. Format:
One line = One prompt in each of the .txt files. Columns: text (the translation), english (the original). Each language has two splits: harmful and harmless.
Sources
All translations are based on the heretic-org/Semantic-Harmless and heretic-org/Semantic-Harmful datasets. For more details, see the README.md file of one of these datasets.
Contributing
You know a language that's not here?
I'd love to add it. You can open a Pull Request for contributing more languages. See `CONTRIBUTING.md` for the details.
License
Copyright © 2026 Vinay Umrethe <umrethevinay@gmail.com>.
This dataset is available under the Creative Commons Attribution Share Alike 4.0 International License.
