Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
Bo Yuan, Danpei Zhao, Zhuoran Liu, Wentao Li, Tian Li
Abstract
Continual learning (CL) breaks off the one-way training manner and enables a model to adapt to new data, semantics and tasks continuously. However, current CL methods mainly focus on single tasks. Besides, CL models are plagued by catastrophic forgetting and semantic drift since the lack of old data, which often occurs in remote-sensing interpretation due to the intricate fine-grained semantics. In this paper, we propose Continual Panoptic Perception (CPP), a unified continual learning model that leverages multi-task joint learning covering pixel-level classification, instance-level segmentation and image-level perception for universal interpretation in remote sensing images. Concretely, we propose a collaborative cross-modal encoder (CCE) to extract the input image features, which supports pixel classification and caption generation synchronously. To inherit the knowledge from the old model without exemplar memory, we propose a task-interactive knowledge distillation (TKD) method, which leverages cross-modal optimization and task-asymmetric pseudo-labeling (TPL) to alleviate catastrophic forgetting. Furthermore, we also propose a joint optimization mechanism to achieve end-to-end multi-modal panoptic perception. Experimental results on the fine-grained panoptic perception dataset validate the effectiveness of the proposed model, and also prove that joint optimization can boost sub-task CL efficiency with over 13% relative improvement on panoptic quality. The project page is available at https://github.com/YBIO/CPP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d688c97-5d00-4671-bf23-b8bd6ae72156Cited by top-tier papers1
Ask how each one uses itBuilds on28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- IL2M: Class Incremental Learning With Dual MemoryEden Belouadah, Adrian PopescuICCV 2019 · 385 citations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
Related papers
- ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt TuningBeomyoung Kim, Joonsang Yu, Sung Ju HwangCVPR 2024
- ADAPT: Attentive Self-Distillation and Dual-Decoder Prediction Fusion for Continual Panoptic SegmentationZe Yang, Shichao Dong, Ruibo Li, Nan Song et al.ICLR 2025
- PLOP: Learning Without Forgetting for Continual Semantic SegmentationArthur Douillard, Yifu Chen, Arnaud Dapogny, Matthieu CordCVPR 2021
- Label-Guided Knowledge Distillation for Continual Semantic Segmentation on 2D Images and 3D Point CloudsZe Yang, Ruibo Li, Evan Ling, Chi Zhang et al.ICCV 2023 · 23 citations
- Class Similarity Weighted Knowledge Distillation for Continual Semantic SegmentationMinh-Hieu Phan, The-Anh Ta, Son Lam Phung, Long Tran-Thanh et al.CVPR 2022 · 57 citations
