Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01figmtu /aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. See our EMNLP 2025 paper for details. tabular1B<n<10B1 likes676 downloads5mo agoHugging Face02figmtu /aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus. This is a subset of the dataset figmtu/aac_c4_deberta_classified. It contains only the sentences that had a dialogue or forum probability of 0.90 or greater. See our EMNLP 2025 paper for details. text100M<n<1B0 likes431 downloads5mo agoHugging Face03hriaz /wikitext-tags-deberta-v31M<n<10M0 likes169 downloads1y agoHugging Face04hlyn-labs /prompt-injection-judge-deberta-datasetgated 🛡️ Prompt Injection Detection Dataset A 400K-sample, production-grade dataset for training binary classifiers to detect prompt injections, jailbreaks, and adversarial attacks targeting LLMs. This is the exact dataset used to train hlyn-labs/prompt-injection-judge-deberta-70m. Quick Start from datasets import load_dataset ds = load_dataset("hlyn-labs/prompt-injection-judge-deberta-dataset") Dataset Summary Stat Value Total Samples 399… See the full description on the dataset page: https://huggingface.co/datasets/hlyn-labs/prompt-injection-judge-deberta-dataset.texttext-classification100K<n<1M4 likes130 downloads5mo agoHugging Face05hriaz /wikitext-tags-deberta-base1M<n<10M0 likes110 downloads1y agoHugging Face06gguichard /wsd_UFSAC_deberta_v3_largetext1M<n<10M0 likes75 downloads2y agoHugging Face07figmtu /aac_c4_deberta_classified_0.90_small_4mThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. This is a subset of the dataset figmtu/aac_c4_deberta_classified. It contains only the sentences that had a dialogue or forum probability of 0.90 or greater. This dataset is further limited to only 4M training examples for use in hyperparameter tuning. See our EMNLP 2025 paper for details. text1M<n<10M0 likes67 downloads5mo agoHugging Face08figmtu /aac_subtitle_deberta_classifiedThis dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. See our EMNLP 2025 paper for details. tabular10M<n<100M0 likes62 downloads5mo agoHugging Face09cike-dev /DeBERTa_multi-class_cb_datasettabular100K<n<1M0 likes60 downloads8mo agoHugging Face10vllm-sr /halueval-spans-deberta HaluEval Span-Level Dataset 🔍 Span-level hallucination detection dataset converted from HaluEval using DeBERTa-FEVER-ANLI NLI model. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans-deberta") Why This Dataset? Problem Solution HaluEval has binary labels only ✅ Span-level annotations Most hallucination datasets are imbalanced ✅ 45.8% hallucinated tokens Token classifiers need… See the full description on the dataset page: https://huggingface.co/datasets/vllm-sr/halueval-spans-deberta.texttoken-classification10K<n<100K0 likes55 downloads9mo agoHugging Face11rosimeirecosta /c_corpus_br_finetuning_language_model_deberta Dataset Card for "c_corpus_br_finetuning_language_model_deberta" More Information needed text100K<n<1M3 likes51 downloads4y agoHugging Face12figmtu /aac_subtitle_deberta_classified_0.75This dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus. This is a subset of the dataset figmtu/aac_subtitle_deberta_classified. It contains only the sentences that had a dialogue or forum probability of 0.75 or greater. See our EMNLP 2025 paper for details. text10M<n<100M0 likes50 downloads5mo agoHugging Face13kkvc-hf /Style-Bert-VITS2-bert_deberta-v2-large-japanese-char-wwm0 likes42 downloads2y agoHugging Face14artianand /test_data_deberta_v3_large_npretabular10K<n<100K0 likes35 downloads2y agoHugging Face15artianand /test_data_deberta_v3_large_racetabular10K<n<100K0 likes31 downloads2y agoHugging Face16Shweta-singh /Deberta_results_racetabular10K<n<100K0 likes30 downloads2y agoHugging Face17artianand /bbq_deberta_v3_large_race_custom_loss_custom_datasettabular10K<n<100K0 likes29 downloads1y agoHugging Face18artianand /bbq_deberta_v3_large_custom_dataset_custom_headtabular10K<n<100K0 likes25 downloads1y agoHugging Face19Shweta-singh /Deberta_results_race_new_input_format_2tabular10K<n<100K0 likes20 downloads2y agoHugging Face20artianand /bbq_deberta_v3_large_race_custom_loss_less_adapter_categories_predictionstabular10K<n<100K0 likes17 downloads2y agoHugging Face21artianand /bbq_deberta_v3_large_race_custom_loss_less_data_predictionstabular10K<n<100K0 likes17 downloads2y agoHugging Face22Crystalcareai /UltraInteract-Debertatext100K<n<1M0 likes15 downloads2y agoHugging Face23artianand /bbq_deberta_v3_large_race_custom_loss_lamda_07_predictionstabular10K<n<100K0 likes15 downloads1y agoHugging Face24AITeamUIT /eval-gliner2-deberta_base-uni-202606210 likes15 downloads4mo agoHugging Face25Shweta-singh /Deberta_results_race_new_input_formattabular10K<n<100K0 likes14 downloads2y agoHugging Face26artianand /deberta_v3_large_race_custom_loss_our_dataset_predictionstabular10K<n<100K0 likes14 downloads2y agoHugging Face27artianand /bbq_deberta_v3_large_race_custom_loss_race_format_predictionstabular10K<n<100K0 likes13 downloads2y agoHugging Face28aimlresearch2023 /climbmix1k-deberta-v3-small1K<n<10K0 likes13 downloads7mo agoHugging Face29RoyArkh /deberta-base-pii-300k100K<n<1M0 likes13 downloads6mo agoHugging Face30AITeamUIT /eval-gliner2-deberta_gliner2_base_fastino-uni-202606220 likes13 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.