datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mms This work presents the most extensive open massively multi-lingual corpus of datasets for training sentiment models.
The corpus consists of 79 manually selected from over 350 datasets reported in the scientific literature based on strict quality criteria and covers 25 languages.
Datasets can be queried using several linguistic and functional features.
In addition, we present a multi-faceted sentiment classification benchmark summarizing hundreds of experiments conducted on different base models, training objectives, dataset collections, and fine-tuning strategies.msb
MSB: Multilingual Safety Benchmark
A red-teaming benchmark of harmful prompts translated into 100 languages,
tested against 6 LLMs, and scored by a 3-judge (Mistral-Large / Qwen3.5 /
Claude Sonnet) consensus for refusal quality and harm.
Provenance
Seed prompts are derived from
nvidia/Aegis-AI-Content-Safety-Dataset-2.0
(CC-BY-4.0). English prompts were deduplicated, human-label-filtered, and
curated down to an 11,788-prompt unsafe pool, then a 10,000-sample… See the full description on the dataset page: https://huggingface.co/datasets/mmse-bench-anon/msb.MM_SFTmmscale-data
MM-SCALE — multimodal moral-acceptability evaluation set
21,977 (image, scenario) pairs across 8,444 DALL·E-generated images, each
labeled with a single mean_rating ∈ [1, 5] for moral acceptability and a
single modality_label ∈ {text, image, both} indicating which modality the
judgment hinges on.
This is a simplified eval view — one row per (image, scenario) pair with one
rating and one modality label. The underlying paired-annotation source (per-
annotator ratings + modality… See the full description on the dataset page: https://huggingface.co/datasets/uunicee/mmscale-data.mmstar_500_lite
mmstar: 500-sample lite benchmark
500 evaluation examples sampled with seed 42. Sampling: uniform random rows without replacement.
Sources and attribution
morpheushoc/MMStar_opencompass, revision 6cc296736fa01334391ec0b819b04d9836e34a3a.
Original dataset and source-image terms apply; this subset does not relicense them. See the linked upstream datasets for their license and citation information.
Reproducibility
subset.json records the pinned… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/mmstar_500_lite.mmscale-data
MM-SCALE — multimodal moral-acceptability evaluation set
21,977 (image, scenario) pairs across 8,444 DALL·E-generated images, each
labeled with a single mean_rating ∈ [1, 5] for moral acceptability and a
single modality_label ∈ {text, image, both} indicating which modality the
judgment hinges on.
This is a simplified eval view — one row per (image, scenario) pair with one
rating and one modality label. The underlying paired-annotation source (per-
annotator ratings + modality votes… See the full description on the dataset page: https://huggingface.co/datasets/mmscale/mmscale-data.MMStar_THMM_SFT_Newhindi-brand24-mms-ipaurdu-brand24-mms-ipamms_yua_resultsmms-fa-alignmentsalgerian-mms-results
