Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
Jinyang Liu, Wondmgezahu Teshome, Sandesh Ghimire, Mario Sznaier, Octavia I. Camps
Abstract
Solving image and video jigsaw puzzles poses the challenging task of rearranging image fragments or video frames from unordered sequences to restore meaningful images and video sequences. Existing approaches often hinge on discriminative models tasked with predicting either the absolute positions of puzzle elements or the permutation actions applied to the original data. Unfortunately, these methods face limitations in effectively solving puzzles with a large number of elements. In this paper, we propose JPDVT, an innovative approach that harnesses diffusion transformers to address this challenge. Specifically, we generate positional information for image patches or video frames, conditioned on their underlying visual content. This information is then employed to accurately assemble the puzzle pieces in their correct positions, even in scenarios involving missing pieces. Our method achieves state-of-the-art performance on several datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac64ee19-7c50-499f-b019-d929de168dcfCited by top-tier papers3
- Generalization Error Analysis for Selective State-Space Models Through the Lens of AttentionArya Honarpisheh, Mustafa Bozdag, Octavia I. Camps, Mario SznaierNeurIPS 2025 · 6 citations
- The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological FragmentsOfir Itzhak Shahar, Gur Elkin, Ohad Ben-ShaharCVPR 2026 · 1 citation
- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosZijia Lu, A S. M. Iftekhar, Gaurav Mittal, Tianjian Meng et al.CVPR 2025
Builds on11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- VDT: General-purpose Video Diffusion Transformers via Mask ModelingHaoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo et al.ICLR 2024 · 117 citations
- DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D ReassemblyGianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio et al.CVPR 2024 · 12 citations
- MultiAnimate: Pose-Guided Image Animation Made ExtensibleYingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An et al.CVPR 2026 · 6 citations
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu et al.AAAI 2024 · 13 citations
- Positional Encoding FieldYunpeng Bai, Haoxiang Li, Qixing HuangICLR 2026 · 22 citations
