Puzzlefusion: Unleashing the Power of Diffusion Models for Spatial Puzzle Solving
Sepidehsadat (Sepid) Hossieni, Mohammad Amin Shabani, Saghar Irandoust, Yasutaka Furukawa
Abstract
This paper presents an end-to-end neural architecture based on Diffusion Models for spatial puzzle solving, particularly jigsaw puzzle and room arrangement tasks. In the latter task, for instance, the proposed system takes a set of room layouts as polygonal curves in the top-down view and aligns the room layout pieces by estimating their 2D translations and rotations, akin to solving the jigsaw puzzle of room layouts. A surprising discovery of the paper is that the simple use of a Diffusion Model effectively solves these challenging spatial puzzle tasks as a conditional generation process. To enable learning of an end-to-end neural system, the paper introduces new datasets with ground-truth arrangements: 1) 2D Voronoi jigsaw dataset, a synthetic one where pieces are generated by Voronoi diagram of 2D pointset; and 2) MagicPlan dataset, a real one offered by MagicPlan from its production pipeline, where pieces are room layouts constructed by augmented reality App by real-estate consumers. The qualitative and quantitative evaluations demonstrate that our approach outperforms the competing methods by significant margins in all the tasks. We will publicly share all our code and data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79bbd633-0765-46de-92e0-1e6b4d2bf3d4Cited by top-tier papers12
- Rectified Point Flow: Generic Point Cloud Pose EstimationTao Sun, Liyuan Zhu, Shengyu Huang, Shuran Song et al.NeurIPS 2025 · 14 citations
- ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco ReconstructionAdeela Islam, Stefano Fiorini, Stuart James, Pietro Morerio et al.ICCV 2025 · 2 citations
- GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular PackingTianyang Xue, Lin Lu, Yang Liu, Mingdong Wu et al.ICCV 2025 · 1 citation
- PuzzleFusion++: Auto-agglomerative 3D Fracture Assembly by Denoise and VerifyZhengqing Wang, Jiacheng Chen, Yasutaka FurukawaICLR 2025
- ShreddingNet: Coarse-to-Fine Restoration for Multi-Source Shredded ManuscriptsHaoyang Cui, Hao Jiang, Yadong MuCVPR 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Deformable Polygonal Flow Matching with Informed Priors and Hierarchical Graph ConstraintsArnaud Gueze, Matthieu Ospici, Damien Rohmer, Marie-Paule CaniAAAI 2026
- DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D ReassemblyGianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio et al.CVPR 2024 · 12 citations
- Solving Masked Jigsaw Puzzles with Diffusion Vision TransformersJinyang Liu, Wondmgezahu Teshome, Sandesh Ghimire, Mario Sznaier et al.CVPR 2024
- Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language ModelsZesen Lyu, Dandan Zhang, Wei Ye, Fangdi Li et al.EMNLP 2025
- HouseDiffusion: Vector Floorplan Generation via a Diffusion Model with Discrete and Continuous DenoisingMohammad Amin Shabani, Sepidehsadat Hosseini, Yasutaka FurukawaCVPR 2023
