VQA-Illusion/MNIST_test
IllusionMNIST — Test Set Dataset summary This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images. The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.
IllusionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images.
The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were produced from these inputs and English scene prompts with a ControlNet variant.
Repository structure
The same filename stem is used across all five directories. For example, <code>Mnist1</code> maps to <code>Mnist1.jpg</code> in every variant directory.
Metadata schema
For either illusionless directory, replace the row's digit target with No illusion.
Label mapping
Download
~~~bash pip install -U huggingface_hub pandas pillow ~~~
~~~python from huggingfacehub import snapshotdownload
datasetdir = snapshotdownload( repoid="VQA-Illusion/MNISTtest", repotype="dataset", ) print(datasetdir) ~~~
Command-line alternative:
~~~bash huggingface-cli download VQA-Illusion/MNISTtest \ --repo-type dataset \ --local-dir MNISTtest ~~~
Load all five image conditions
~~~python from pathlib import Path import pandas as pd from huggingfacehub import snapshotdownload
root = Path(snapshotdownload( repoid="VQA-Illusion/MNISTtest", repotype="dataset", )) df = pd.readcsv(root / "dfdata.csv", dtype={"label": "int64"})
folders = { "illusion": "illimages", "illusionfiltered": "illusionimagesfiltered", "illusionless": "illusionlessimages", "illusionlessfiltered": "illusionlessimagesfiltered", "raw": "rawimages", } for condition, folder in folders.items(): df[condition + "path"] = df["image_name"].map( lambda name, folder=folder: root / folder / (name + ".jpg") )
records = [] for row in df.itertuples(index=False): for condition in folders: labelid = 10 if condition.startswith("illusionless") else int(row.label) records.append({ "imagename": row.imagename, "condition": condition, "imagepath": getattr(row, condition + "path"), "labelid": labelid, "labeltext": "No illusion" if labelid == 10 else "digit " + str(labelid), })
evaluationdf = pd.DataFrame(records) assert evaluationdf["image_path"].map(Path.exists).all() ~~~
Filtered variants
The released filtered images were produced with the preprocessing evaluated in the paper: Gaussian, averaging, and median blurs followed by grayscale conversion and sharpening. Appendix K provides the exact OpenCV implementation and parameters.
Intended use and evaluation
Use this split for zero-shot or fine-tuned digit-illusion recognition, VQA, No illusion rejection, and paired robustness comparisons. Recommended classification metrics are accuracy, macro precision, macro recall, and macro F1. Keep all conditions for the same <code>image_name</code> together when creating any derived partitions.
Dataset creation and safety
The authors generated English scene descriptions with several language models and used a ControlNet variant to combine them with resized MNIST source-condition images. Human reviewers validated quality. The paper reports that the public release was screened with NSFW detectors and that flagged images were excluded.
The paper reports 1,219 IllusionMNIST test samples, while the current repository has 1,109 rows in <code>df_data.csv</code>. This card describes the files currently hosted; use the current metadata file for reproducible indexing.
Important usage notes
- Treat <code>df_data.csv</code> as the authoritative index.
- Hugging Face may auto-detect top-level folders as <code>imagefolder</code> classes. These are image conditions, not digit targets.
- The correct target for both illusionless variants is No illusion (ID 10), irrespective of <code>label</code>.
- The benchmark primarily contains one large hidden category per image; consult the paper for full limitations.
License
This dataset repository declares the MIT license. Users should also review and comply with any applicable terms associated with MNIST and other upstream components.
Citation
~~~bibtex @misc{rostamkhani2024illusoryvqa, title = {Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions}, author = {Rostamkhani, Mohammadmostafa and Ansari, Baktash and Sabzevari, Hoorieh and Rahmani, Farzan and Eetemadi, Sauleh}, year = {2024}, eprint = {2412.08169}, archivePrefix = {arXiv}, primaryClass = {cs.CV}, url = {https://arxiv.org/abs/2412.08169} } ~~~
Contact
Questions and reproducibility issues can be submitted through the official GitHub repository.
