Team Ai
Datasetpublic

HexQuant/vdr-multilingual-train

Multilingual Visual Document Retrieval Dataset This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B). It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model. How it was created This is the entire data pipeline used to create the Italian subset of this dataset. Each… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/vdr-multilingual-train.

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
0likes156downloads
settings

This repository belongs to HexQuant on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namevdr-multilingual-train
visibilitypublic
licenceapache-2.0
gatedno
ownerHexQuant
Account settings
HexQuant/vdr-multilingual-train · Team Ai