Team Ai
Datasetpublic

VQA-Illusion/MNIST_test

IllusionMNIST — Test Set Dataset summary This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images. The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes5.3kdownloads
Dataset Card

IllusionMNIST — Test Set

Dataset summary

This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images.

The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were produced from these inputs and English scene prompts with a ControlNet variant.

PropertyValue
Hugging Face repositoryVQA-Illusion/MNIST_test
Official splitTest
TaskIllusion digit classification / visual question answering
Annotated base examples1,109
Image variants per example5
Image formatJPEG
Metadata file<code>df_data.csv</code>
PaperarXiv:2412.08169
CodeIllusoryVQA/IllusoryVQA

Repository structure

PathFilesDescriptionEvaluation target
<code>ill_images/</code>1,109Generated images containing a digit illusion.Digit in <code>label</code>.
<code>illusionimagesfiltered/</code>1,109Illusion images processed with the paper's filter pipeline.Digit in <code>label</code>.
<code>illusionless_images/</code>1,109Matched scene images without an embedded illusion.No illusion.
<code>illusionlessimagesfiltered/</code>1,109Filtered illusionless controls.No illusion.
<code>raw_images/</code>1,109MNIST source-condition images used to guide generation.Digit in <code>label</code>.
<code>df_data.csv</code>1Canonical metadata for the base examples.—
<code>captions.csv</code>1Pool of 1,027 English scene descriptions used in generation.—

The same filename stem is used across all five directories. For example, <code>Mnist1</code> maps to <code>Mnist1.jpg</code> in every variant directory.

Metadata schema

ColumnTypeDescription
<code>image_name</code>stringImage identifier and shared filename stem.
<code>Pprompt</code>stringPositive scene prompt used during generation.
<code>Nprompt</code>stringNegative generation prompt; currently <code>low quality</code>.
<code>illusion_strength</code>floatControl strength used during illusion generation; currently <code>1.5</code>.
<code>label</code>integer-like stringGround-truth digit, from <code>0</code> to <code>9</code>.

For either illusionless directory, replace the row's digit target with No illusion.

Label mapping

Numeric IDClass labelStored test value
0digit 0<code>0</code>
1digit 1<code>1</code>
2digit 2<code>2</code>
3digit 3<code>3</code>
4digit 4<code>4</code>
5digit 5<code>5</code>
6digit 6<code>6</code>
7digit 7<code>7</code>
8digit 8<code>8</code>
9digit 9<code>9</code>
10No illusionDerived target for the two illusionless directories

Download

~~~bash pip install -U huggingface_hub pandas pillow ~~~

~~~python from huggingfacehub import snapshotdownload

datasetdir = snapshotdownload( repoid="VQA-Illusion/MNISTtest", repotype="dataset", ) print(datasetdir) ~~~

Command-line alternative:

~~~bash huggingface-cli download VQA-Illusion/MNISTtest \ --repo-type dataset \ --local-dir MNISTtest ~~~

Load all five image conditions

~~~python from pathlib import Path import pandas as pd from huggingfacehub import snapshotdownload

root = Path(snapshotdownload( repoid="VQA-Illusion/MNISTtest", repotype="dataset", )) df = pd.readcsv(root / "dfdata.csv", dtype={"label": "int64"})

folders = { "illusion": "illimages", "illusionfiltered": "illusionimagesfiltered", "illusionless": "illusionlessimages", "illusionlessfiltered": "illusionlessimagesfiltered", "raw": "rawimages", } for condition, folder in folders.items(): df[condition + "path"] = df["image_name"].map( lambda name, folder=folder: root / folder / (name + ".jpg") )

records = [] for row in df.itertuples(index=False): for condition in folders: labelid = 10 if condition.startswith("illusionless") else int(row.label) records.append({ "imagename": row.imagename, "condition": condition, "imagepath": getattr(row, condition + "path"), "labelid": labelid, "labeltext": "No illusion" if labelid == 10 else "digit " + str(labelid), })

evaluationdf = pd.DataFrame(records) assert evaluationdf["image_path"].map(Path.exists).all() ~~~

Filtered variants

The released filtered images were produced with the preprocessing evaluated in the paper: Gaussian, averaging, and median blurs followed by grayscale conversion and sharpening. Appendix K provides the exact OpenCV implementation and parameters.

Intended use and evaluation

Use this split for zero-shot or fine-tuned digit-illusion recognition, VQA, No illusion rejection, and paired robustness comparisons. Recommended classification metrics are accuracy, macro precision, macro recall, and macro F1. Keep all conditions for the same <code>image_name</code> together when creating any derived partitions.

Dataset creation and safety

The authors generated English scene descriptions with several language models and used a ControlNet variant to combine them with resized MNIST source-condition images. Human reviewers validated quality. The paper reports that the public release was screened with NSFW detectors and that flagged images were excluded.

The paper reports 1,219 IllusionMNIST test samples, while the current repository has 1,109 rows in <code>df_data.csv</code>. This card describes the files currently hosted; use the current metadata file for reproducible indexing.

Important usage notes

  • —Treat <code>df_data.csv</code> as the authoritative index.
  • —Hugging Face may auto-detect top-level folders as <code>imagefolder</code> classes. These are image conditions, not digit targets.
  • —The correct target for both illusionless variants is No illusion (ID 10), irrespective of <code>label</code>.
  • —The benchmark primarily contains one large hidden category per image; consult the paper for full limitations.

License

This dataset repository declares the MIT license. Users should also review and comply with any applicable terms associated with MNIST and other upstream components.

Citation

~~~bibtex @misc{rostamkhani2024illusoryvqa, title = {Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions}, author = {Rostamkhani, Mohammadmostafa and Ansari, Baktash and Sabzevari, Hoorieh and Rahmani, Farzan and Eetemadi, Sauleh}, year = {2024}, eprint = {2412.08169}, archivePrefix = {arXiv}, primaryClass = {cs.CV}, url = {https://arxiv.org/abs/2412.08169} } ~~~

Contact

Questions and reproducibility issues can be submitted through the official GitHub repository.