tasksource/lewidi
lewidi Learning with Disagreements (LeWiDi, SemEval-2023 Task 11 and its 2025 edition): soft labels from every annotator. One config per dataset: md_agreement (offensiveness, 5 annotators), hs_brexit (hate speech, 6), armis (Arabic misogyny and sexism, 3), conv_abuse (abuse in user turns of chatbot dialogues, 3 or more), csc (sarcasm rated 1-6), mp (MultiPICo irony, multilingual) and varierrnli (NLI where each annotator may accept several labels). soft_label lists the share of… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/lewidi.
lewidi
Learning with Disagreements (LeWiDi, SemEval-2023 Task 11 and its 2025 edition): soft labels from every annotator.
One config per dataset: mdagreement (offensiveness, 5 annotators), hsbrexit (hate speech, 6), armis (Arabic misogyny and sexism, 3), convabuse (abuse in user turns of chatbot dialogues, 3 or more), csc (sarcasm rated 1-6), mp (MultiPICo irony, multilingual) and varierrnli (NLI where each annotator may accept several labels). ``softlabel` lists the share of annotators per label in labels` order; varierrnli instead gives, per NLI label, the share of annotators who accepted it. Text fields are unpacked from the harmonized JSON. The Paraphrase set (an 11-level scale over 500 items) is left out. The original datasets' licenses and terms apply unchanged; see the LeWiDi repository (https://github.com/Le-Wi-Di/le-wi-di.github.io) and each dataset's paper.
Repackaged as parquet for tasksource by scripts/upload_repackaged.py.
