datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vdr-multilingual-train
Multilingual Visual Document Retrieval Dataset
This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B).
It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model.
How it was created
This is the entire data pipeline used to create the Italian subset of this dataset. Each step… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-train.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.multilingual-vlm-reasoning
Multilingual VLM Visual Reasoning Benchmark
Paper: Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
Overview
This dataset contains ~1,000 visual reasoning questions translated from English into 6 Indian languages:
Hindi (hi), Tamil (ta), Telugu (te), Bengali (bn), Kannada (kn), Marathi (mr).
Source benchmarks: MathVista (testmini), ScienceQA, MMMU
Translation: IndicTrans2 (AI4Bharat), verified against GPT-4o/Gemini… See the full description on the dataset page: https://huggingface.co/datasets/Swastikr/multilingual-vlm-reasoning.vdr-multilingual-train-corpusMMc-Instruct-Stage2DrivingVQA-conflictvdr-multilingual-trainmultilingual-coco
Multilingual Common Objects in Context (COCO) Dataset
This dataset is a collection of multiple language open-source captions of COCO dataset.
The split in this dataset is set according to Andrej Karpathy's split from dataset_coco.json file. The collection was created specifically for simplicity of use in training and evaluation pipeline by non-commercial and research purposes. The COCO images dataset is licensed under a Creative Commons Attribution 4.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/multilingual-coco.counterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/counterfactual-pendulum-multilingual.Synthdog-Multilingual-100
Synthdog Multilingual
The Synthdog dataset created for training in Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model.
Using the official Synthdog code, we created >1 million training samples for improving OCR capabilities in Large Vision-Language Models.
Dataset Details
We provide the images for download in two .tar.gz files. Download and extract them in folders of the same name (so cat images.tar.gz.* | tar xvzf -C images; tar xvzf… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/Synthdog-Multilingual-100.multilingual-llava-bench-in-the-wild
🌍 PALO: A Polyglot Large Multimodal Model for 5B People
Vision-language conversation in English, Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali and Urdu.
Multi-lingual Evaluation Dataset
This repository contains LLaVA Bench In-the-Wild, translated to Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali, and Urdu.
Please refer to our paper for details.
madove-multilingual-train
MaDOVE Train Split
This dataset extends llamaindex/vdr-multilingual-train for visual question answering (VQA).
Short and long answers are generated to the questions provided in the original dataset using Qwen2.5-VL-72B-Instruct.
Statistics
Language
Short Answer
Long Answer
German
44,580
47,990
English
42,158
46,303
Spanish
47,733
45,261
French
42,049
45,558
Italian
43,109
45,316
TOTAL
219,629
230,428
License
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/m-elio/madove-multilingual-train.sea-vl-conflict
SEA-VL Conflict
Image-text conflict dataset built from
SEACrowd/sea-vl_crowdsourcing
(English captions of culturally-relevant Southeast Asian images).
Conflicts were hand-authored: for each caption, one object or attribute (a dish ingredient,
a landmark's location, a color, a count, etc.) is flipped to a different but plausible value.
Each example pairs an image with a truthful original_caption and a conflicting_caption
that alters exactly one object or attribute, creating an… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/sea-vl-conflict.multilingual_rksvdr-multilingual-train
Multilingual Visual Document Retrieval Dataset
This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B).
It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model.
How it was created
This is the entire data pipeline used to create the Italian subset of this dataset. Each step… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/vdr-multilingual-train.worldcuisines-conflict
WorldCuisines Conflict
Image-text conflict dataset built from
worldcuisines/vqa (task1, English prompts).
The true dish name (the VQA answer) is the image_bias; the conflicting caption asserts a
wrong multiple-choice option (text_bias), and a second wrong option is the distractor.
The question is the dataset's English open-ended prompt.
Each example pairs an image with a truthful original_caption and a conflicting_caption
that alters exactly one object or attribute, creating an… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/worldcuisines-conflict.remote_sensing_VQA_multilingual
Remote Sensing VQA — Multilingual
A multilingual counterfactual MCQ dataset built from remote sensing / satellite imagery.
Each row contains a satellite image, two captions (original vs counterfactual), and a multiple-choice question probing whether a VLM follows the image or the misleading text.
Languages
Language
Code
Rows
English
en
50
Hindi
hi
50
Urdu
ur
50
Telugu
te
50
Bahasa Indonesia
id
50
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/remote_sensing_VQA_multilingual.indian-crafts-captions-multilingualmultilingual_ocr_llm_2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/multilingual_ocr_llm_2.flickr30k-multilingual
Flickr 30k Multilingual
Multilingual version of the Flickr30k dataset. This dataset contains images paired with original English captions, along with synthetic translations into Polish, German, French, and Spanish generated using Meta's NLLB-200-1.3B model.
🚀 Quickstart
You can easily load this dataset using the Hugging Face datasets library:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("aplominski/flickr30k-multilingual")
#… See the full description on the dataset page: https://huggingface.co/datasets/aplominski/flickr30k-multilingual.github-readme-retrieval-multilingual
GitHub Readme Retrieval
This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual.github-readme-retrieval-multilingual_deprecated
GitHub Readme Retrieval
This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_deprecated.pendulum-conflict
Pendulum Conflict Dataset
This dataset is a manually curated subset of 100 isolated multimodal conflict examples derived from the akomand/counterfactual_pendulum dataset.
It is designed for evaluating Multimodal Large Language Models (MLLMs) under controlled visual-textual conflicts.
Dataset Statistics
Total Samples: 100
Categories: angle (25), light (25), shadow_len (25), shadow_pos (25)
Language: English
Corrected / Audited Samples: 20
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/pendulum-conflict.xfund-multilingualindic-multilingual-image-captions
Indic Multilingual Image Caption Dataset
This dataset contains 4,500 unique images with captions in:
English
Hindi
Bengali
Tamil
Source composition
3,000 images from COCO Caption 2017
1,500 images from TextCaps
Each image is stored once and paired with four multilingual caption
variants. Hindi, Bengali and Tamil captions were generated from the
selected English captions using the NLLB-200 distilled translation model.
Intended use
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/arikatokachi/indic-multilingual-image-captions.magazines-multilingual-vqa
Magazines Multilingual VQA
A multilingual Visual Question Answering dataset built from 29,039 public-domain magazine and newspaper pages sourced from archive.org. Each page has:
Verbatim OCR in the page's native language
English description, page type, scan-quality metadata
One grounded VQA pair in one of 10 target languages (round-robin assigned)
Full provenance and license carried from the source archive.org record
Annotations generated with Google Gemma 4 31B via vLLM.… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/magazines-multilingual-vqa.rpg-conflict
RPG Fantasy Battle Conflict Dataset
A dataset of 100 visual RPG combat conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between battle screenshots and caption text). Derived from the rcannizzaro/rpg_fantasy_battle_counterfactual_v2 dataset.
Dataset Statistics
This dataset consists of a single train split containing 100 perfectly isolated conflict samples derived from the… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/rpg-conflict.vdr-multilingual-test
Multilingual Visual Document Retrieval Benchmarks
This dataset consists of 15 different benchmarks used to initially evaluate the vdr-2b-multi-v1 multimodal retrieval embedding model. These benchmarks allow the testing of multilingual, multimodal retrieval capabilities on text-only, visual-only and mixed page screenshots.
Each language subset contains queries and images in that language and is divided into three different categories by the "pagetype" column. Each category contains… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-test.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.coco-counterfactual-conflict
COCO-Counterfactual Conflict
Image-text conflict dataset built from
Intel/COCO-Counterfactuals.
Each COCO-Counterfactuals example is a minimal pair of captions differing by a single noun
subject, with a matching image for each. We keep the truthful image (image_0) and its
caption as original_caption, and use the counterfactual caption as conflicting_caption.
The swapped noun is extracted automatically (image_bias = true noun, text_bias = altered
noun); the question and… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/coco-counterfactual-conflict.
