pradhaansbhat/Thinking-In-Boxes
[NeurIPS-2026] Thinking in Boxes: 3D Editing in Real Images Made Easy
<div> <img src="assets/thumbnail.png"> </div>
**Pradhaan S Bhat**<sup>1</sup><sup>∗</sup> · **Naveen Chandra R**<sup>1</sup><sup>∗</sup> · **Rishubh Parihar**<sup>1</sup> · **Vaibhav Vavilala**<sup>2</sup> · **R. Venkatesh Babu**<sup>1</sup> · **D.A. Forsyth** · **Anand Bhattad**<sup>4</sup>
<sup>1</sup> Indian Institute of Science <sup>2</sup> Apple <sup>3</sup> UIUC <sup>4</sup> Johns Hopkins University
<sup>∗</sup> Equal Contribution
    ![Demo]()
<div> <img src="assets/teaser.png"> </div>
This is the trained LoRA used in the paper Thinking In Boxes: 3D Editing in Real Images Made Easy.
Thinking-In-Boxes is an Image-to-Image Generative LoRA for black-forest-labs/FLUX.1-Kontext-dev trained for the task of Geometric Image Editing.
This LoRA is trained on the Thinking-In-Boxes dataset available here. Details of training are available in the supplementary section of the paper.
Usage with Diffusers 🧨
Please refer to the GitHub Repository on setup, inference and training.
Citation
If you find our work useful, please consider citing:
@misc{bhat2026thinkingboxes3dediting,
title = {Thinking in Boxes: 3D Editing in Real Images Made Easy},
author = {Pradhaan S Bhat and Naveen Chandra R and Rishubh Parihar and Vaibhav Vavilala and R. Venkatesh Babu and D. A. Forsyth and Anand Bhattad},
year = {2026},
eprint = {2606.20556},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2606.20556}
}