shahab7899/Hyperphantasia
A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs Mohammad Shahab Sepehri Berk Tinaz Zalan Fabian Mahdi Soltanolkotabi Github Repository Hyperphantasia is a synthetic Visual Question Answering (VQA) benchmark dataset that probes the mental visualization capabilities of Multimodal Large Language Models (MLLMs) from a vision perspective. We reveal that state-of-the-art models struggle with simple tasks that require visual… See the full description on the dataset page: https://huggingface.co/datasets/shahab7899/Hyperphantasia.
<p align="center"> <img src="logo.png" alt="drawing" width="700" style="float: center;"/> </p> <!-- <h1 align="center">Hyperphantasia</h1> --> <h1 align="center">A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs</h1> <!-- <h3 align="center">Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi</h3> -->
<p align="center"> <a href="https://scholar.google.com/citations?user=j2scUKoAAAAJ&hl=en">Mohammad Shahab Sepehri</a> <a href="https://scholar.google.com/citations?user=5EKjsXQAAAAJ&hl=en">Berk Tinaz</a> <a href="https://scholar.google.com/citations?user=5EKjsXQAAAAJ&hl=en">Zalan Fabian</a> <a href="https://scholar.google.com/citations?user=narJyMAAAAAJ&hl=en">Mahdi Soltanolkotabi</a> </p>
<p align="center"> <a href="https://github.com/AIF4S/Hyperphantasia">Github Repository</a> </p>
<p align="center"> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg" alt="License"/></a> </p>
<p align="justify" > Hyperphantasia is a synthetic Visual Question Answering (VQA) benchmark dataset that probes the mental visualization capabilities of Multimodal Large Language Models (MLLMs) from a vision perspective. We reveal that state-of-the-art models struggle with simple tasks that require visual simulation and imagination. Our dataset consists of 1200 samples with four different puzzles in two categories of Interpolation and Extrapolation. </p> <p align="center"> <img src="overview.png" alt="drawing" width="550" style="float: center;"/> </p> <p align="justify"> Hyperphantasia has three levels of difficulty to evaluate the extent and generalizability of mental visualization capabilities of MLLMs. </p>
<p align="center"> <img src="difficulities.png" alt="drawing" width="550" style="float: center;"/> </p> <p align="justify">
Usage
You can find our evaluation code in our Github repository.
Acknowledgement
We would like to thank Microsoft for an Accelerating Foundation Models Research grant that provided the OpenAI credits enabling this work. This research was also in part supported by AWS credits through an Amazon Faculty Research Award and a NAIRR Pilot Award. M. Soltanolkotabi and MS. Sepehri were supported by the USC–Capital One Center for Responsible AI and Decision Making in Finance (CREDIF) Fellowship. M. Soltanolkotabi is also supported by the Packard Fellowship in Science and Engineering, a Sloan Research Fellowship in Mathematics, an NSF-CAREER under award \#1846369, DARPA FastNICS program, and NSF-CIF awards \#1813877 and \#2008443. and NIH DP2LM014564-01.
