Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dbarbedillo /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K14 likes643 downloads4y agoHugging Face02Kenpache /multilingual-financial-sentiment Multilingual Financial Sentiment Dataset A curated dataset of 39,829 financial news sentences annotated with sentiment labels (Negative / Neutral / Positive) across 7 languages, collected from 80+ financial news sources worldwide. Dataset Summary Total samples 39,829 Languages 7 (EN, ZH, JA, DE, FR, ES, AR) Labels 3 (negative, neutral, positive) Format CSV Sources 80+ financial news outlets Languages Language Code Samples %… See the full description on the dataset page: https://huggingface.co/datasets/Kenpache/multilingual-financial-sentiment.texttext-classification10K<n<100K0 likes435 downloads6mo agoHugging Face03Tulsiandhare /Multilingual_medical_symptom_triage tags: - medical - healthcare - classification - outbreak-detection - triage - multilingual - adaption - india Multilingual Medical Symptom Triage Dataset Dataset Description A Mutlilingual medical triage dataset containing 9,064 patient symptom descriptions in Hindi, English, and Hinglish (code-mixed Hindi-English), paired with triage recommendations and rich clinical metadata. Designed for training multilingual triage classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.text10K<n<100K1 likes339 downloads3mo agoHugging Face04ameyhengle /Multilingual-Needle-in-a-Haystack Multilingual Needle in a Haystack (MLNeedle) The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.text10K<n<100K3 likes309 downloads1y agoHugging Face05JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes304 downloads6mo agoHugging Face06KumarSahil299885 /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K3 likes288 downloads3mo agoHugging Face07Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes263 downloads2y agoHugging Face08FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes229 downloads2y agoHugging Face09vanila434 /multilingual-elder-safety-msgs multilingual-elder-safety-msgs A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation. Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.texttext-classification1K<n<10K0 likes200 downloads6mo agoHugging Face10Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes175 downloads2y agoHugging Face11iNLP-Lab /multilingual-lima Multilingual LIMA A multilingual extension of the LIMA instruction-tuning dataset. The original English prompt–response pairs were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt User instruction (translated; en is the original). output Assistant response (translated; en is the original). Languages (configs): en (original), zh, it, bn… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-lima.texttext-generation10K<n<100K0 likes172 downloads5mo agoHugging Face12danielelvs /multilingual-islr-mediapipe Multilingual ISLR MediaPipe Landmarks Dataset Description This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning. The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.tabularvideo-classification1K<n<10K0 likes149 downloads1d agoHugging Face13iNLP-Lab /multilingual-s1 Multilingual s1 A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.texttext-generation1K<n<10K0 likes124 downloads5mo agoHugging Face14flax-community /conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.text10M<n<100M1 likes108 downloads3y agoHugging Face15AnasAlokla /multilingual_go_emotions Overview: This dataset is updated from on the go_emotions dataset With the same labels, but add 5 new languages: Arabic, French, Spanish , Dutch, and Turkish. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English Arabic, French, Spanish , Dutch, and Turkish texttext-classification100K<n<1M1 likes106 downloads1y agoHugging Face16flax-community /conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following: train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.text10M<n<100M0 likes94 downloads3y agoHugging Face17FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes92 downloads5mo agoHugging Face18gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes90 downloads2y agoHugging Face19BeitTigreAI /tigre-data-parallel-multilingual Tigre Parallel Corpus This corpus has 66,144 sentence and phrase pairs. Every pair has Tigre (ትግሬ, Ethi_tig) as the source. The targets are Tigrinya, English, Arabic, Swedish and German. The pairs were gathered from three sources: SMOL, a community-contributed file and Tatoeba. Each source was cleaned and put into the same five-column format. Tigre is a Semitic language spoken mainly in Eritrea and written in Ge'ez (Ethiopic) script. Very little parallel data exists for it.… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.texttranslation100K<n<1M1 likes89 downloads2d agoHugging Face20iNLP-Lab /multilingual-safety Multilingual Safety Instructions A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt Harmful user instruction (translated; en is the original). output Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.texttext-generation10K<n<100K0 likes89 downloads5mo agoHugging Face21Rapidata /multilingual-llm-jokes-4o-claude-gemini Rapidata Generated Joke Preference Dataset We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'. It took us less than 5 days to get all of the responses. The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.tabular1K<n<10K14 likes86 downloads1y agoHugging Face22flax-community /conceptual-12m-mbart-50-multilingualimage10M<n<100M2 likes85 downloads5y agoHugging Face23Steveeeeeeen /multilingual_evalstabularn<1K0 likes85 downloads4mo agoHugging Face24Debk /Indian-Multilingual-Bias-Dataset Indian Multilingual Bias Dataset Dataset Description The Indian Multilingual Bias Dataset is a comprehensive collection designed to evaluate and measure social biases in Large Language Models (LLMs) across three major Indian languages: English, Bengali (বাংলা), and Hindi (हिंदी). This dataset is based on the original Indian-BhED dataset and focuses on four critical dimensions of bias prevalent in Indian society. Key Features 🌐 Multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Indian-Multilingual-Bias-Dataset.texttext-classification1K<n<10K0 likes80 downloads3mo agoHugging Face25Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes77 downloads28d agoHugging Face26erickfmm /agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs The code for processing can be found here Useful for data distillation, training or benchmarking. Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.tabularsentence-similarity1M<n<10M0 likes74 downloads1y agoHugging Face27flax-community /multilingual-vqatext1M<n<10M0 likes63 downloads3y agoHugging Face28joshdavham /multilingual-frequency-lists Multilingual Frequency Lists This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese. Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'. These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.tabular10K<n<100K1 likes62 downloads5mo agoHugging Face29molamin /Kinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English. The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples. text100K<n<1M0 likes60 downloads4y agoHugging Face30karanverma19 /Indian_Multilingual_Scam_Message_Dataset Indian Multilingual Scam Message Dataset Overview This dataset contains 120 realistic SMS and text messages from Indian contexts, labeled as scam or legitimate. It reflects real-world communication patterns across Hindi, Hinglish, and English. Features 120 high-quality samples Multilingual (Hindi, Hinglish, English) Real-world inspired scam and legitimate messages Includes reasoning for each label Covers multiple domains: banking, ecommerce, telecom, utilities… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Indian_Multilingual_Scam_Message_Dataset.textn<1K1 likes59 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.