deberta
Datasets
All datasets matching “deberta”aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
See our EMNLP 2025 paper for details.
prompt-injection-judge-deberta-dataset
🛡️ Prompt Injection Detection Dataset
A 400K-sample, production-grade dataset for training binary classifiers to detect prompt injections, jailbreaks, and adversarial attacks targeting LLMs.
This is the exact dataset used to train hlyn-labs/prompt-injection-judge-deberta-70m.
Quick Start
from datasets import load_dataset
ds = load_dataset("hlyn-labs/prompt-injection-judge-deberta-dataset")
Dataset Summary
Stat
Value
Total Samples
399… See the full description on the dataset page: https://huggingface.co/datasets/hlyn-labs/prompt-injection-judge-deberta-dataset.wikitext-tags-deberta-basewsd_UFSAC_deberta_v3_largeaac_c4_deberta_classified_0.90_small_4mThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
This dataset is further limited to only 4M training examples for use in hyperparameter tuning.
See our EMNLP 2025 paper for details.
