Team Ai
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face02saurabh5 /code_rlvr_mixture_dpotabular10K<n<100K0 likes28 downloads11mo agoHugging Face03sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes20 downloads2y agoHugging Face04mlfoundations-dev /hero_run_3_code_mix_all_shardstabular100K<n<1M0 likes20 downloads1y agoHugging Face05taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes16 downloads1y agoHugging Face06watchstep /ko-en-code-mixing-sts Korean–English Code-Mixing STS Dataset This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred. Interactive Dashboard 🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/ The dashboard provides: Interactive data exploration and filtering Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.tabularsentence-similarity1K<n<10K0 likes16 downloads1y agoHugging Face07atx-labs /marathi-codemix-qagated Marathi Minglish QA ~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles. Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms. Example Question: Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi? Answer: Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.tabulartext-generation1M<n<10M1 likes13 downloads2mo agoHugging Face08hamishivi /code_rlvr_mixture_dpotabular10K<n<100K0 likes10 downloads1y agoHugging Face09hamishivi /code_rlvr_mixture_sfttabular10K<n<100K0 likes8 downloads1y agoHugging Face10mlfoundations-dev /train_fasttext_classifier_seed_code_worst_mix_12_8_5_3_1tabularn<1K0 likes5 downloads2y agoHugging Face11mlfoundations-dev /train_fasttext_classifier_seed_code_best_mixtabularn<1K0 likes5 downloads2y agoHugging Face12mlfoundations-dev /train_fasttext_classifier_seed_code_worst_mix_3_1tabularn<1K0 likes4 downloads2y agoHugging Face13mlfoundations-dev /train_fasttext_classifier_seed_code_worst_mix_8_5_3_1tabularn<1K0 likes4 downloads2y agoHugging Face14nlp-dataset /Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset Sinhala-English Hospitality Review Corpus This dataset contains 4,133 user reviews collected from Facebook hotel pages across Sri Lanka, annotated at the sentence level for sentiment analysis. The reviews come from 400 hotels across 144 tourist destinations, including Colombo, Kandy, Anuradhapura, Negombo, and other major regions. Annotation Details All samples were manually annotated by the author to ensure consistency in sentiment labeling. We welcome volunteers… See the full description on the dataset page: https://huggingface.co/datasets/nlp-dataset/Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset.tabular1K<n<10K0 likes4 downloads14h agoHugging Face15mlfoundations-dev /train_fasttext_classifier_seed_code_worst_mix_5_3_1tabularn<1K0 likes3 downloads2y agoHugging Face16teppap /code-mix-thai-engtabular1K<n<10K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.