Team Ai
20 results

deberta

figmtu /aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. See our EMNLP 2025 paper for details. tabular1B<n<10B1 likes681 downloads5mo agoHugging Facefigmtu /aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus. This is a subset of the dataset figmtu/aac_c4_deberta_classified. It contains only the sentences that had a dialogue or forum probability of 0.90 or greater. See our EMNLP 2025 paper for details. text100M<n<1B0 likes430 downloads5mo agoHugging Facehlyn-labs /prompt-injection-judge-deberta-datasetgated 🛡️ Prompt Injection Detection Dataset A 400K-sample, production-grade dataset for training binary classifiers to detect prompt injections, jailbreaks, and adversarial attacks targeting LLMs. This is the exact dataset used to train hlyn-labs/prompt-injection-judge-deberta-70m. Quick Start from datasets import load_dataset ds = load_dataset("hlyn-labs/prompt-injection-judge-deberta-dataset") Dataset Summary Stat Value Total Samples 399… See the full description on the dataset page: https://huggingface.co/datasets/hlyn-labs/prompt-injection-judge-deberta-dataset.texttext-classification100K<n<1M4 likes121 downloads5mo agoHugging Facehriaz /wikitext-tags-deberta-base1M<n<10M0 likes113 downloads1y agoHugging Facegguichard /wsd_UFSAC_deberta_v3_largetext1M<n<10M0 likes77 downloads2y agoHugging Facefigmtu /aac_c4_deberta_classified_0.90_small_4mThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. This is a subset of the dataset figmtu/aac_c4_deberta_classified. It contains only the sentences that had a dialogue or forum probability of 0.90 or greater. This dataset is further limited to only 4M training examples for use in hyperparameter tuning. See our EMNLP 2025 paper for details. text1M<n<10M0 likes67 downloads5mo agoHugging Face