Two-stage Visual Cues Enhancement Network for Referring Image Segmentation
Yang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen, Yu-Gang Jiang, Xiaolin Wei, Lin Ma
Abstract
Referring Image Segmentation (RIS) aims at segmenting the target object from an image referred by one given natural language expression. The diverse and flexible expressions as well as complex visual contents in the images raise the RIS model with higher demands for investigating fine-grained matching behaviors between words in expressions and objects presented in images. However, such matching behaviors are hard to be learned and captured when the visual cues of referents (i.e. referred objects) are insufficient, as the referents with weak visual cues tend to be easily confused by cluttered background at boundary or even overwhelmed by salient objects in the image. And the insufficient visual cues issue can not be handled by the cross-modal fusion mechanisms as done in previous work. In this paper, we tackle this problem from a novel perspective of enhancing the visual information for the referents by devising a Two-stage Visual cues enhancement Network (TV-Net), where a novel Retrieval and Enrichment Scheme (RES) and an Adaptive Multi-resolution feature Fusion (AMF) module are proposed. Specifically, RES retrieves the most relevant image from an external data pool with regard to both the visual and textual similarities, and then enriches the visual information of the referent with the retrieved image for better multimodal feature learning. AMF further enhances the visual detailed information by incorporating the highresolution feature maps from lower convolution layers of the image. Through the two-stage enhancement, our proposed TV-Net enjoys better performances in learning fine-grained matching behaviors between the natural language expression and image, especially when the visual information of the referent is inadequate, thus produces better segmentation results. Extensive experiments are conducted to validate the effectiveness of the proposed method on the RIS task, with our proposed TV-Net surpassing the state-of-theart approaches on four benchmark datasets. Our code is available at: https://github.com/SxJyJay/TV-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3334efd6-a152-4096-add0-ba534cbc69faCited by top-tier papers11
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 48 citations
- Temporal Collection and Distribution for Referring Video Object SegmentationJiajin Tang, Ge Zheng, Sibei YangICCV 2023 · 44 citations
- Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal ModelsYang Jiao, Shaoxiang Chen, Zequn Jie, Jingjing Chen et al.NeurIPS 2024 · 27 citations
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang et al.ACM MM 2022 · 12 citations
- Suspected Objects Matter: Rethinking Model's Prediction for One-stage Visual GroundingYang Jiao, Zequn Jie, Jingjing Chen, Lin Ma et al.ACM MM 2023 · 7 citations
Builds on6
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- See-Through-Text Grouping for Referring Image SegmentationDing-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen et al.ICCV 2019 · 153 citations
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu et al.ACM MM 2020 · 74 citations
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Bi-Directional Relationship Inferring Network for Referring Image SegmentationZhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang et al.CVPR 2020
Related papers
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie et al.ACM MM 2023 · 34 citations
- Encoder Fusion Network With Co-Attention Embedding for Referring Image SegmentationGuang Feng, Zhiwei Hu, Lihe Zhang, Huchuan LuCVPR 2021
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang et al.CVPR 2021
- Bottom-Up Shift and Reasoning for Referring Image SegmentationSibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou et al.CVPR 2021
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu et al.CVPR 2024
