datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gujarati-english-codemixed-sentiment
Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset
Dataset Summary
This dataset addresses an empirically confirmed gap in the code-mixed
sentiment analysis literature: a systematic literature review of 48
retrieved records (39 core studies) found zero existing
Gujarati-English code-mixed sentiment datasets, despite Gujarati having
over 55 million native speakers. This dataset provides the first
Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.code_rlvr_mixture_dpoCode_Mixed_video_Complainthero_run_3_code_mix_all_shardsArzEn-CodeMixed
ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset
This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants.
Dataset Details
Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.ko-en-code-mixing-sts
Korean–English Code-Mixing STS Dataset
This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred.
Interactive Dashboard
🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/
The dashboard provides:
Interactive data exploration and filtering
Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.marathi-codemix-qa
Marathi Minglish QA
~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles.
Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms.
Example
Question:
Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi?
Answer:
Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.code_rlvr_mixture_dpocode_rlvr_mixture_sfttrain_fasttext_classifier_seed_code_worst_mix_12_8_5_3_1train_fasttext_classifier_seed_code_best_mixtrain_fasttext_classifier_seed_code_worst_mix_3_1train_fasttext_classifier_seed_code_worst_mix_8_5_3_1Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset
Sinhala-English Hospitality Review Corpus
This dataset contains 4,133 user reviews collected from Facebook hotel pages across Sri Lanka, annotated at the sentence level for sentiment analysis. The reviews come from 400 hotels across 144 tourist destinations, including Colombo, Kandy, Anuradhapura, Negombo, and other major regions.
Annotation Details
All samples were manually annotated by the author to ensure consistency in sentiment labeling. We welcome volunteers… See the full description on the dataset page: https://huggingface.co/datasets/nlp-dataset/Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset.train_fasttext_classifier_seed_code_worst_mix_5_3_1code-mix-thai-eng
