datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
DeBERTa_multi-class_cb_datasetbbq_deberta_v3_large_race_custom_loss_custom_datasetbbq_deberta_v3_large_custom_dataset_custom_headtest_data_deberta_v3_large_racetest_data_deberta_v3_large_nprebbq_deberta_v3_large_race_custom_loss_less_data_predictionsaac_subtitle_deberta_classifiedThis dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
bbq_deberta_v3_large_race_custom_loss_less_adapter_categories_predictionsDeberta_results_race_new_input_format_2bbq_deberta_v3_large_race_custom_loss_race_format_predictionsDeberta_results_racedeberta_v3_large_race_custom_loss_our_dataset_predictionsbbq_deberta_v3_large_race_custom_loss_lamda_07_predictionsbbq_deberta_v3_large_5_categories_finetuned_predictionsDeberta_results_race_new_input_formattest_data_deberta_v3_large_racebbq_deberta_v3_large_race_custom_loss_predictionsbbq_deberta_v3_large_race_custom_loss_single_adapter_predictionsDeBERTa_CB_hp_tuning_splitdeberta_rmdeberta_v3_large_race_custom_loss_fusion_our_dataset_predictionsbbq_deberta_v3_large_race_finetuned_predictionspersonalization_prompt_response_oasst_deberta_v3deberta_rm_completebbq_deberta_v3_large_race_custom_loss_changed_adapter_categories_predictionsbbq_deberta_v3_large_race_custom_loss_custom_dataset_custom_headbbq_deberta_v3_large_race_custom_loss_our_datasetpersonalization_prompt_response_oasst_deberta_v3_oldresults_joke_gen_mistral_dpo_ze_zeggen_dat_deberta_test
