HexQuant/vdr-multilingual-train
Multilingual Visual Document Retrieval Dataset This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B). It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model. How it was created This is the entire data pipeline used to create the Italian subset of this dataset. Each… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/vdr-multilingual-train.
This repository belongs to HexQuant on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
