datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmless_alpacaMultilingual-Harmless-Harmful
Multilingual Harmless and Harmful Prompts
What is this?
This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 9 languages (including original English data).
This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Multilingual-Harmless-Harmful.harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
Semantic-Harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.Multilingual-Harmless-Harmful
Multilingual Harmless and Harmful Prompts
What is this?
This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 8 languages (including original English data).
This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/Heyjab/Multilingual-Harmless-Harmful.OSI-BenchGenerated_Injected_PDFs_HARMLESS
Generated Injected PDFs — HARMLESS
A synthetic dataset of 1,100 PDF files built for training and evaluating structural PDF-malware detectors. It pairs benign PDFs with PDFs into which safe, non-executable "malware-shaped" objects have been injected, so a model can learn to separate the two from byte-level structure alone.
⚠️ Safety notice — read first
Nothing in this dataset is real malware. Every injected payload is built from industry-standard, non-executable… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS.cai-conversation-harmless
Dataset Card for "cai-conversation-dev1705629166"
More Information needed
meta-llama_Llama-3.1-8B-Instruct-jdgfct-Harmlessnessharmful_harmless_instructions
Dataset Card for "harmful_harmless_instructions"
More Information needed
hh-harmless-base-qwen3-8b-margin-dpo-margin-logsQwen_Qwen2-7B-Instruct-jdgfct-HarmlessnessNexusflow_Athene-70B-jdgfct-Harmlessnesstrojan-harmless-rlhf-goldenHARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.grok-conversation-harmless
Dataset Card for "cai-conversation-dev1705950597"
More Information needed
grok-conversation-harmless-old
Dataset Card for "cai-conversation-dev1705369037"
More Information needed
grok-conversation-harmless2
Dataset Card for "cai-conversation-dev1705680551"
More Information needed
harmless-aira-dataset
Harmless-Aira Dataset
Dataset Summary
This dataset contains a collection of prompt + completion examples of LLM following instructions in a conversational manner. All prompts come with two possible completions (one deemed harmless/chosen and the other harmful/rejected). The dataset is available in both Portuguese and English.
Supported Tasks and Leaderboards
This dataset can be utilized to train a reward/preference model or DPO fine-tuning.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/harmless-aira-dataset.dpo-q2572b-a70b-jllm3-Harmlessness-Areward-bench-hacking-rewards-harmless-train-normaldpo-qwen2572b-llama3170b-jdg-Llama3-HarmlessnessAnthropic-harmless-baseharmless_alpaca
[!NOTE]
This is a "Just in case" mirror of mlabonne/harmless_alpaca
hh-rlhf-harmless-base-rollouts-gpt-oss-20b-diverse-openrouterharmless-poisoned-0.04-SUDO-murderRiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards.
The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers.
Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy.
The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.anthropic-helpful-harmless-rlhfcai-conversation-harmless-old
Dataset Card for "cai-conversation-harmless"
More Information needed
harmful-harmless-prompts-library
Harmful vs Harmless Prompts Dataset
This dataset aggregates multiple sources of jailbreak prompts, harmful queries, and benign prompts.
Labels
0 = harmless / regular prompt
2 = harmful or jailbreak-related prompt
Sources
Includes datasets from:
TrustAIRLab in-the-wild jailbreak prompts (MIT License)
DiegoAI597 harmful actions [Apache 2.0]
djapp18 JailbreaksOverTime [CC-BY-4.0]
Bravansky compact jailbreaks [MIT]
Splits
train… See the full description on the dataset page: https://huggingface.co/datasets/aplominski/harmful-harmless-prompts-library.
