ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps
Xingke Song, Xiaoying Yang, Chenglin Yao, Jianfeng Ren, Ruibin Bai, Xin Chen, Xudong Jiang
Abstract
Solving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To tackle these challenges, we propose a framework of Evolutionary Reinforcement Learning with Multi-head Puzzle Perception (ERL-MPP) to derive a better set of swapping actions for solving the puzzles. Specifically, to tackle the challenges of perceiving the puzzle with gaps, a Multi-head Puzzle Perception Network (MPPN) with a shared encoder is designed, where multiple puzzlet heads comprehensively perceive the local assembly status, and a discriminator head provides a global assessment of the puzzle. To explore the large swapping action space efficiently, an Evolutionary Reinforcement Learning (EvoRL) agent is designed, where an actor recommends a set of suitable swapping actions from a large action space based on the perceived puzzle status, a critic updates the actor using the estimated rewards and the puzzle status, and an evaluator coupled with evolutionary strategies evolves the actions aligning with the historical assembly experience. The proposed ERL-MPP is comprehensively evaluated on the JPLEG-5 dataset with large gaps and the MIT dataset with large-scale puzzles. It significantly outperforms all state-of-the-art models on both datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e2a85ba-44d7-47e6-93d6-9e11080cc84fCited by top-tier papers1
Ask how each one uses itBuilds on4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Fine-Grained Object Classification via Self-Supervised Pose AlignmentXuhui Yang, Yaowei Wang, Ke Chen, Yong Xu et al.CVPR 2022 · 91 citations
- Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision TransformersBin Ren, Yahui Liu, Yue Song, Wei Bi et al.CVPR 2023
- Solving Jigsaw Puzzles With Eroded BoundariesDov Bridger, Dov Danon, Ayellet TalCVPR 2020
Related papers
- Siamese-Discriminant Deep Reinforcement Learning for Solving Jigsaw Puzzles with Large Eroded GapsXingke Song, Jiahuan Jin, Chenglin Yao, Shihe Wang et al.AAAI 2023 · 24 citations
- CEARI: Co-Evolutionary Agents for Reassembling and Inpainting Puzzles with Gaps and Missing PiecesXingke Song, Jianxu Shangguan, Yiran Li, Jialu Zhang et al.ACM MM 2025 · 1 citation
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language ModelsYu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao et al.ICLR 2026 · 11 citations
- Brick-by-Brick: Combinatorial Construction with Deep Reinforcement LearningHyunsoo Chung, Jungtaek Kim, Boris Knyazev, Jinhwi Lee et al.NeurIPS 2021 · 29 citations
- Visual Jigsaw Post-Training Improves MLLMsPenghao Wu, Yushan Zhang, Haiwen Diao, Bo Li et al.ICLR 2026 · 25 citations
