bryanzhou008/mini-vlm-toolkit-data
mini-vlm-toolkit counting & grounding benchmarks Processed evaluation splits used by mini-vlm-toolkit. Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running python run_benchmark.py --data… See the full description on the dataset page: https://huggingface.co/datasets/bryanzhou008/mini-vlm-toolkit-data.
mini-vlm-toolkit counting & grounding benchmarks
Processed evaluation splits used by mini-vlm-toolkit. Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation.
You normally don't need to download anything by hand: running python run_benchmark.py --data <split> ... fetches the TSV, the images and (for ODinW) the GT files into $LMUData on first use.
Sources, licenses and citations
This repo only re-packages existing public benchmarks; all credit goes to the original authors, and use is subject to the original licenses/terms.
- PixmoCount — allenai/pixmo-count (ODC-BY). Images were downloaded from their original Flickr URLs, which are subject to their individual licenses; rows whose image could not be downloaded are dropped, so this is the fixed subset we evaluate on. Deitke et al., Molmo and PixMo, 2024.
- CountQA — Jayant-Sravan/CountQA (CC BY 4.0).
- FSC147 — Ranjan et al., Learning To Count Everything, CVPR 2021 (code/data).
- FSCD-147 / FSCD-LVIS — Nguyen et al., Few-shot Object Counting and Detection, ECCV 2022 (code/data). FSCD-LVIS images are from COCO / LVIS (Gupta et al., CVPR 2019) and follow their terms of use.
- ODinW-13 — Li et al., Grounded Language-Image Pre-training (GLIP), CVPR 2022, and Li et al., ELEVATER, NeurIPS 2022. Domains originate from Roboflow public datasets and PASCAL VOC and follow their respective licenses.
If you are an author of one of these datasets and would like it removed, please open a discussion on this repo.
