Team Ai
Datasetpublic

md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset

Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 3y agoView on Hugging Face
1likes39downloads
3 commits on main
81dd1973y ago

Update README.md

md-nishat-008
143c57b3y ago

Upload 3 files

md-nishat-008
51e10ea3y ago

initial commit

Md Nishat Raihan