datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private.multi-model-cot-missing-answers
Multi-model CoT responses without a final answer
Teacher responses from the text part of the next_jev stage 1 data (built from JonesLin/multi-model-cot-2730)
whose final_answer is empty, collected so they can be re-run with a model.
One row per response: 23,615 train and 500 validation rows.
The answer_and_cot stage 1 objective trains the backbone to emit the final answer, so it needs an answer
for every response. These rows have none.
Why the answer is missing… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/multi-model-cot-missing-answers.lm-eval-results-MTSAIR-multi_verse_model-private
Dataset Card for Evaluation run of MTSAIR/multi_verse_model
Dataset automatically created during the evaluation run of model MTSAIR/multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MTSAIR-multi_verse_model-private.Multimodel_Redteaming_Data
🛡️ Multimodal Redteaming (EN, FR, DE, IT, ES)
A high-quality multilingual red teaming dataset designed to evaluate the robustness and safety of Large Language Models (LLMs) against adversarial prompts. The dataset includes both text-only and image-supported conversations with expert-curated annotations for AI safety evaluation, benchmarking, and alignment research.
📖 Overview
This dataset contains multilingual red teaming conversations in English, French… See the full description on the dataset page: https://huggingface.co/datasets/Nawras-99/Multimodel_Redteaming_Data.rhdem-multimodel-200spq-2026-06-22rhdem-multimodel-2026-06-22_160429
