Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes261 downloads2y agoHugging Face02FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes256 downloads2y agoHugging Face03Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes178 downloads2y agoHugging Face04danielelvs /multilingual-islr-mediapipe Multilingual ISLR MediaPipe Landmarks Dataset Description This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning. The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.tabularvideo-classification1K<n<10K0 likes177 downloads6d agoHugging Face05gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes119 downloads2y agoHugging Face06Rapidata /multilingual-llm-jokes-4o-claude-gemini Rapidata Generated Joke Preference Dataset We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'. It took us less than 5 days to get all of the responses. The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.tabular1K<n<10K14 likes91 downloads1y agoHugging Face07erickfmm /agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs The code for processing can be found here Useful for data distillation, training or benchmarking. Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.tabularsentence-similarity1M<n<10M0 likes77 downloads1y agoHugging Face08Steveeeeeeen /multilingual_evalstabularn<1K0 likes75 downloads4mo agoHugging Face09Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes60 downloads1mo agoHugging Face10wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes51 downloads1mo agoHugging Face11fufu1976 /multilingual-shopping-queries Multilingual Shopping Queries → Marketplace Queries 941 pairs mapping how shoppers actually phrase a product search in their own language to the short English query that marketplace listings (AliExpress) are indexed under. It comes from OneFindMe, a free multilingual AI product search engine for AliExpress (voice, image and plain-language search in 12 languages). Why this is useful Cross-lingual product search does not fail on vocabulary alone. A shopper's… See the full description on the dataset page: https://huggingface.co/datasets/fufu1976/multilingual-shopping-queries.tabulartranslationn<1K0 likes51 downloads17d agoHugging Face12freococo /quran_multilingual_parallel 📘 Qur’an Multilingual Parallel Dataset (quran_multilingual_parallel) This dataset presents a clean, structurally-aligned multilingual parallel corpus of the Qur’anic text. It is intended for linguistic, computational, and cross-lingual AI applications — not only for religious interpretation. It contains over 6,200 verse-level alignments in 54 human languages, formatted in a machine-friendly .csv structure with language-specific translation fields. 🧠 Dataset Highlights… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran_multilingual_parallel.tabulartranslation1K<n<10K5 likes45 downloads1y agoHugging Face13Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes42 downloads5mo agoHugging Face14joshdavham /multilingual-frequency-lists Multilingual Frequency Lists This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese. Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'. These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.tabular10K<n<100K1 likes38 downloads5mo agoHugging Face15hjm1980 /korea-places-multilingual Korean Place Names, Multilingual Built and maintained by Korea Basics, a sourced guide to Korean entry rules and getting around, published in seven languages. 16,126 places in South Korea with their Korean (Hangul) name next to the romanized English name, plus Japanese and Chinese names where the source has them, coordinates, road-name address, and subway lines for stations. Why this exists A visitor who reads "Gyeongbokgung Palace" in a guide cannot type that… See the full description on the dataset page: https://huggingface.co/datasets/hjm1980/korea-places-multilingual.tabular10K<n<100K1 likes27 downloads2mo agoHugging Face16helinivan /sarcasm_headlines_multilingual Dataset Card for Multilingual Sarcasm Detection Dataset Summary Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source. The sources of news articles are: The Huffington Post (en, non-sarcastic) The Onion (en, sarcastic) NOS (nl, non-sarcastic) De Speld (nl, sarcastic) Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.tabular10K<n<100K1 likes26 downloads4y agoHugging Face17PalakEngineerMaster /Processed_TTS_Multilingual_Data Processed TTS Multilingual Data Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages. Datasets Included Subset Samples Hours Description indic_voices_r 239,684 548.8h Indic Voices_R — IVR recordings rasa 201,509 361.2h RASA — read speech (wiki, conv, book, news) indictts_iitm 155,236 253.6h Indic TTS (IIT Madras) — studio TTS recordings at 48kHz Total 596,429 1,163.6h Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.tabulartext-to-speech100K<n<1M0 likes25 downloads8mo agoHugging Face18jamesdborin /Nemotron-SFT-Multilingual-v1-prompt-only Nemotron-SFT-Multilingual-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v1-prompt-only.tabular1M<n<10M0 likes21 downloads3mo agoHugging Face19jamesdborin /Nemotron-SFT-Multilingual-v2-prompt-only Nemotron-SFT-Multilingual-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v2-prompt-only.tabular100K<n<1M0 likes21 downloads3mo agoHugging Face20wow2000 /multilingual_jailbreak_challengesgatedtabular1K<n<10K2 likes20 downloads2y agoHugging Face21FrancophonIA /multilingualcrowspairs [!NOTE] Dataset origin: https://gitlab.inria.fr/corpus4ethics/multilingualcrowspairs/ MultiLingualCrowsPairs Multilingual CrowS-Pairs, a challenge dataset for measuring stereotypical biases present in the masked language models (MLMs) in 7 different languages. This challenge dataset was built on the Crows-Pairs corpus (Nangia et al. 2020) using the methodology described in (Névéol et al. 2023). The 7 new languages are the following: Arabic from Maghreb and the Arab world in… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingualcrowspairs.tabulartext-classification10K<n<100K1 likes18 downloads2y agoHugging Face22aq1048576 /red_team_agent_analysis_multilingual_story_analysis_detailed red_team_agent_analysis_multilingual_story_analysis_detailed This dataset was automatically uploaded from the red-team-agent repository. Dataset Information Original file: multilingual_story_analysis_detailed.csv Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis/multilingual_story_analysis_detailed.csv Validation: Valid CSV with 1000 rows, 12 columns (0.1MB) Usage import pandas as pd from datasets import load_dataset # Load using datasets… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_multilingual_story_analysis_detailed.tabularother1K<n<10K0 likes15 downloads1y agoHugging Face23Dvvreddy /multilingual_abusive-non-abusivetabular100K<n<1M1 likes10 downloads4mo agoHugging Face24Sowmya15 /gibberish_multilingualtabular10K<n<100K1 likes9 downloads3y agoHugging Face25Shoriful025 /Multilingual_E-commercetabularn<1K0 likes6 downloads10mo agoHugging Face26Tonic /multilingual-EMIR-reporting-csvtabular10K<n<100K0 likes6 downloads2mo agoHugging Face27model2me /scienceqa-multilingual-hindi#ScienceQA Hindi Translation Dataset ##Dataset Description This dataset is a Hindi-translated version of the original ScienceQA dataset. It includes multiple-choice science questions, with fields for: Images (optional visual context), Hints (optional support text), English questions and their Hindi translations, Multiple answer choices, Correct answers. This translation is intended to support multilingual education research, question-answering in Hindi, and fairness studies in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/model2me/scienceqa-multilingual-hindi.tabular10K<n<100K0 likes5 downloads1y agoHugging Face28infinite-dataset-hub /MultilingualTranscriptionDataset MultilingualTranscriptionDataset tags: Transcription, LanguageProcessing, Multilingual Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'MultilingualTranscriptionDataset' is a curated collection of text transcriptions from various audio recordings. Each transcription is provided in multiple languages, emphasizing the diversity and complexity of language processing. This dataset aims to assist in developing machine learning… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MultilingualTranscriptionDataset.tabularn<1K0 likes4 downloads2y agoHugging Face29MithuSi /multi-lingual-llm Dataset Card for Dataset Name This data set contains set questions in tamil and possible answers, with the correct answer in the column. This helps to test LLM to see for accuracy. Dataset Details Dataset Description Curated by: Madhumitha Sivalingapandian Language(s) (NLP): Tamil and English License: [More Information Needed] Uses Used for measuring performance of LLM. Direct Use Spoken language accuracy measurement dataset… See the full description on the dataset page: https://huggingface.co/datasets/MithuSi/multi-lingual-llm.tabularn<1K0 likes2 downloads2y agoHugging Face30isc-tleavitt /opus100-multilingualtabular1K<n<10K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.