Team Ai
Datasetpublic

md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset

Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 3y agoView on Hugging Face
1likes39downloads

md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset · main · files are served by the source, never re-hosted here