Team Ai
Datasetpublic

bryanzhou008/mini-vlm-toolkit-data

mini-vlm-toolkit counting & grounding benchmarks Processed evaluation splits used by mini-vlm-toolkit. Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation. You normally don't need to download anything by hand: running python run_benchmark.py --data… See the full description on the dataset page: https://huggingface.co/datasets/bryanzhou008/mini-vlm-toolkit-data.

sourceHugging Faceotherupdated 12d agoView on Hugging Face
0likes43downloads
Dataset Card

mini-vlm-toolkit counting & grounding benchmarks

Processed evaluation splits used by mini-vlm-toolkit. Each <split>.tsv holds the exact prompts (question) and annotations (answer plus metadata) we evaluate on; image_path is relative to images/<folder>/ inside images/<folder>.zip. GLIP.zip holds the ODinW-13 configs and COCO-format val/test annotations used for ODinW AP evaluation.

You normally don't need to download anything by hand: running python run_benchmark.py --data <split> ... fetches the TSV, the images and (for ODinW) the GT files into $LMUData on first use.

SplitRowsImage archive
PixmoCount_test527images/PixmoCount.zip
PixmoCount_val534images/PixmoCount.zip
CountQA1528images/CountQA.zip
FSC147_test1190images/FSC147.zip
FSC147_val1286images/FSC147.zip
FSCD147bbox_test1190images/FSC147.zip
FSCD147bbox_val1286images/FSC147.zip
FSCDLVISbbox_test1014images/LVIS.zip
FSCDLVISbbox_val1181images/LVIS.zip
ODinW13bboxtest_qwen34575images/ODinW.zip
ODinW13bboxval_qwen34938images/ODinW.zip

Sources, licenses and citations

This repo only re-packages existing public benchmarks; all credit goes to the original authors, and use is subject to the original licenses/terms.

  • —PixmoCount — allenai/pixmo-count (ODC-BY). Images were downloaded from their original Flickr URLs, which are subject to their individual licenses; rows whose image could not be downloaded are dropped, so this is the fixed subset we evaluate on. Deitke et al., Molmo and PixMo, 2024.
  • —CountQA — Jayant-Sravan/CountQA (CC BY 4.0).
  • —FSC147 — Ranjan et al., Learning To Count Everything, CVPR 2021 (code/data).
  • —FSCD-147 / FSCD-LVIS — Nguyen et al., Few-shot Object Counting and Detection, ECCV 2022 (code/data). FSCD-LVIS images are from COCO / LVIS (Gupta et al., CVPR 2019) and follow their terms of use.
  • —ODinW-13 — Li et al., Grounded Language-Image Pre-training (GLIP), CVPR 2022, and Li et al., ELEVATER, NeurIPS 2022. Domains originate from Roboflow public datasets and PASCAL VOC and follow their respective licenses.

If you are an author of one of these datasets and would like it removed, please open a discussion on this repo.