SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection
Jia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong, Guodao Zhang, Zhi Liu, Jiyong Zhang
Abstract
Recently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have attempted to directly apply the foundation model to the RGB-D video salient object detection (RGB-D VSOD) task, which often encounters three challenges, including the dependence on manual prompts, the high memory consumption of sequential adapters, and the computational burden of memory attention. To address the limitations, we propose a novel method, namely Segment Anything Model with Depth-guided Adaptive Queries (SAM-DAQ), which adapts SAM2 to pop-out salient objects from videos by seamlessly integrating depth and temporal cues within a unified framework. Firstly, we deploy a parallel adapter-based multi-modal image encoder (PAMIE), which incorporates several depth-guided parallel adapters (DPAs) in a skip-connection way. Remarkably, we fine-tune the frozen SAM encoder under prompt-free conditions, where the DPA utilizes depth cues to facilitate the fusion of multi-modal features. Secondly, we deploy a query-driven temporal memory (QTM) module, which unifies the memory bank and prompt embeddings into a learnable pipeline. Concretely, by leveraging both frame-level queries and video-level queries simultaneously, the QTM module can not only selectively extract temporal consistency features but also iteratively update the temporal representations of the queries. Extensive experiments are conducted on three RGB-D VSOD datasets, and the results show that the proposed SAM-DAQ consistently outperforms state-of-the-art methods in terms of all evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8033c24c-8f27-4662-b23b-c880465eca09Builds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei et al.ICML 2023 · 388 citations
- Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object DetectionShixuan Gao, Pingping Zhang, Tianyu Yan, Huchuan LuACM MM 2024 · 93 citations
- Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything ModelZihan Zhong, Zhiqiang Tang, Tong He, Haoyang Fang et al.ICLR 2024 · 91 citations
Related papers
- M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object DetectionJiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong et al.CVPR 2026
- Improving SAM for Camouflaged Object Detection via Dual Stream AdaptersJiaming Liu, Linghe Kong, Guihai ChenICCV 2025 · 5 citations
- Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object DetectionWenjun Hui, Zhenfeng Zhu, Shuai Zheng, Yao ZhaoCVPR 2024
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
- Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object DetectionZhenni Yu, Xiaoqin Zhang, Li Zhao, Yi Bin et al.ACM MM 2024 · 41 citations
