Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlabonne /harmless_alpacatext10K<n<100K49 likes24k downloads2y agoHugging Face02heretic-org /Multilingual-Harmless-Harmful Multilingual Harmless and Harmful Prompts What is this? This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 9 languages (including original English data). This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Multilingual-Harmless-Harmful.text1K<n<10K4 likes1.3k downloads5d agoHugging Face03OS-Software /harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF text10K<n<100K0 likes1k downloads4mo agoHugging Face04heretic-org /Semantic-Harmless [!IMPORTANT] You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful Semantic Harmful-Harmless Prompt Pairs Summary This dataset contains one-to-one semantic matches between prompts from two source datasets: mlabonne/harmful_behaviors mlabonne/harmless_alpaca The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.textn<1K4 likes736 downloads4mo agoHugging Face05Heyjab /Multilingual-Harmless-Harmful Multilingual Harmless and Harmful Prompts What is this? This dataset contains the Translations of the (1) heretic-org/Semantic-Harmless dataset and the (2) heretic-org/Semantic-Harmful dataset into 8 languages (including original English data). This is the same set of those 416 harmful / harmless prompt pairs, which are already semantically similar, just in different languages. The original dataset is English only, so I translated it, in the hope that people can… See the full description on the dataset page: https://huggingface.co/datasets/Heyjab/Multilingual-Harmless-Harmful.text1K<n<10K1 likes685 downloads18d agoHugging Face06HarmlessSR07 /OSI-Benchtabularvisual-question-answering1K<n<10K4 likes647 downloads10mo agoHugging Face07Cyber-security-final-project /Generated_Injected_PDFs_HARMLESS Generated Injected PDFs — HARMLESS A synthetic dataset of 1,100 PDF files built for training and evaluating structural PDF-malware detectors. It pairs benign PDFs with PDFs into which safe, non-executable "malware-shaped" objects have been injected, so a model can learn to separate the two from byte-level structure alone. ⚠️ Safety notice — read first Nothing in this dataset is real malware. Every injected payload is built from industry-standard, non-executable… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS.documenttabular-classification1K<n<10K0 likes297 downloads3mo agoHugging Face08HuggingFaceH4 /cai-conversation-harmless Dataset Card for "cai-conversation-dev1705629166" More Information needed text10K<n<100K17 likes276 downloads3y agoHugging Face09penfever /meta-llama_Llama-3.1-8B-Instruct-jdgfct-Harmlessnesstext100K<n<1M0 likes249 downloads6mo agoHugging Face10longphann /harmful_harmless_instructions Dataset Card for "harmful_harmless_instructions" More Information needed textn<1K4 likes234 downloads3y agoHugging Face11W-61 /hh-harmless-base-qwen3-8b-margin-dpo-margin-logstabular1K<n<10K0 likes226 downloads7mo agoHugging Face12penfever /Qwen_Qwen2-7B-Instruct-jdgfct-Harmlessnesstext100K<n<1M0 likes212 downloads6mo agoHugging Face13penfever /Nexusflow_Athene-70B-jdgfct-Harmlessnesstext100K<n<1M0 likes204 downloads2y agoHugging Face14Baidicoot /trojan-harmless-rlhf-goldentext10K<n<100K0 likes172 downloads2y agoHugging Face15Cyber-security-final-project /HARMLESS_Synthetic_Injected_PDFs_EDA Injected PDFs - EDA and Evaluation Corpus This repository holds the exploratory data analysis for a project on detecting harmless-but-real attack payloads injected into PDF files, together with the dataset that analysis produced. The project has two halves, both in the notebook Final_project_V7_EDA.ipynb: Question Input Part 1 Is our synthetic corpus a stand-in for real malware, or is it something else? The published CIC feature table (11,126 x 34) Part 2 Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.imagetext-classification1K<n<10K0 likes166 downloads2mo agoHugging Face16HuggingFaceH4 /grok-conversation-harmless Dataset Card for "cai-conversation-dev1705950597" More Information needed 29 likes126 downloads1y agoHugging Face17HuggingFaceH4 /grok-conversation-harmless-old Dataset Card for "cai-conversation-dev1705369037" More Information needed text10K<n<100K1 likes117 downloads3y agoHugging Face18HuggingFaceH4 /grok-conversation-harmless2 Dataset Card for "cai-conversation-dev1705680551" More Information needed text10K<n<100K9 likes102 downloads3y agoHugging Face19nicholasKluge /harmless-aira-dataset Harmless-Aira Dataset Dataset Summary This dataset contains a collection of prompt + completion examples of LLM following instructions in a conversational manner. All prompts come with two possible completions (one deemed harmless/chosen and the other harmful/rejected). The dataset is available in both Portuguese and English. Supported Tasks and Leaderboards This dataset can be utilized to train a reward/preference model or DPO fine-tuning. Languages… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/harmless-aira-dataset.texttext-classification10K<n<100K6 likes84 downloads1y agoHugging Face20penfever /dpo-q2572b-a70b-jllm3-Harmlessness-Atext100K<n<1M0 likes82 downloads2y agoHugging Face21Ayush-Singh /reward-bench-hacking-rewards-harmless-train-normaltabular1K<n<10K0 likes72 downloads2y agoHugging Face22penfever /dpo-qwen2572b-llama3170b-jdg-Llama3-Harmlessnesstext100K<n<1M0 likes63 downloads2y agoHugging Face23thobauma /Anthropic-harmless-basetext10K<n<100K0 likes61 downloads2y agoHugging Face24heretic-org /harmless_alpaca [!NOTE] This is a "Just in case" mirror of mlabonne/harmless_alpaca text10K<n<100K4 likes61 downloads5mo agoHugging Face25MWilinski /hh-rlhf-harmless-base-rollouts-gpt-oss-20b-diverse-openroutertextn<1K0 likes51 downloads7mo agoHugging Face26thobauma /harmless-poisoned-0.04-SUDO-murdertext10K<n<100K1 likes50 downloads3y agoHugging Face27Ray2333 /RiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards. The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers. Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy. The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.tabular100K<n<1M0 likes48 downloads2y agoHugging Face28Baidicoot /anthropic-helpful-harmless-rlhftext100K<n<1M0 likes46 downloads2y agoHugging Face29HuggingFaceH4 /cai-conversation-harmless-old Dataset Card for "cai-conversation-harmless" More Information needed text10K<n<100K3 likes45 downloads3y agoHugging Face30aplominski /harmful-harmless-prompts-library Harmful vs Harmless Prompts Dataset This dataset aggregates multiple sources of jailbreak prompts, harmful queries, and benign prompts. Labels 0 = harmless / regular prompt 2 = harmful or jailbreak-related prompt Sources Includes datasets from: TrustAIRLab in-the-wild jailbreak prompts (MIT License) DiegoAI597 harmful actions [Apache 2.0] djapp18 JailbreaksOverTime [CC-BY-4.0] Bravansky compact jailbreaks [MIT] Splits train… See the full description on the dataset page: https://huggingface.co/datasets/aplominski/harmful-harmless-prompts-library.texttext-classification10K<n<100K0 likes45 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.