yurkes/patch_tasks_vllm
Dataset Card for Patch-Based Visual Question Answering Dataset Dataset Details Dataset Description This dataset contains approximately 305,000 triplets of question, answer, and image designed for patch-based visual reasoning tasks. A standard question in this dataset is formatted as follows: Image Grid: The image is divided into a 4x4 grid of 16 equal-sized patches. Patches are numbered sequentially from the top-left corner and moving right, then… See the full description on the dataset page: https://huggingface.co/datasets/yurkes/patch_tasks_vllm.
338
