Guided Slot Attention for Unsupervised Video Object Segmentation
Minhyeok Lee, Suhwan Cho, Dogyoon Lee, Chaewon Park, Jungho Lee, Sangyoun Lee
Abstract
Unsupervised video object segmentation aims to segment the most prominent object in a video sequence. However, the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue, we propose a guided slot attention network to reinforce spatial structural information and obtain better foreground-background separation. The foreground and background slots, which are initialized with query guidance, are iteratively refined based on interactions with template information. Furthermore, to improve slot-template interaction and effectively fuse global and local features in the target and reference frames, K-nearest neighbors filtering and a feature aggregation transformer are introduced. The proposed model achieves state-of-the-art performance on two popular datasets. Additionally, we demonstrate the robustness of the proposed model in challenging scenes through various comparative experiments. Code and models are available at https://github.com/ Hydragon516/GSANet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfa60ac0-75e0-450c-b20e-c972bbd7a4c4Cited by top-tier papers7
- Object-Centric Refinement for Enhanced Zero-Shot SegmentationSrinivasa Rao Nandam, Sara Atito Ali, Zhenhua Feng, Josef Kittler et al.ICLR 2026 · 5 citations
- Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric LearningWonJun Moon, Hyun Seok Seong, Jae-Pil HeoCVPR 2026 · 3 citations
- RoMo: Robust Motion Segmentation Improves Structure from MotionLily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker et al.ICCV 2025 · 2 citations
- Factor-Wise Homogeneity of Slot-Attention for Continual Object-Centric LearningIlmin Kang, Hoyong Kim, Seungju Bang, Minwoo Kang et al.ICML 2026
- Segment Any Motion in VideosNan Huang, Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer et al.CVPR 2025
Builds on15
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object DetectionHongwei Zhu, Peng Li, Haoran Xie, Xuefeng Yan et al.AAAI 2022 · 242 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
Related papers
- Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in VideosRui Qian, Shuangrui Ding, Xian Liu, Dahua LinICCV 2023 · 23 citations
- Dual Prototype Attention for Unsupervised Video Object SegmentationSuhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee et al.CVPR 2024
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 29 citations
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
