Two-stage Visual Cues Enhancement Network for Referring Image Segmentation
Yang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen, Yu-Gang Jiang, Xiaolin Wei, Lin Ma
摘要
Referring Image Segmentation (RIS) aims at segmenting the target object from an image referred by one given natural language expression. The diverse and flexible expressions as well as complex visual contents in the images raise the RIS model with higher demands for investigating fine-grained matching behaviors between words in expressions and objects presented in images. However, such matching behaviors are hard to be learned and captured when the visual cues of referents (i.e. referred objects) are insufficient, as the referents with weak visual cues tend to be easily confused by cluttered background at boundary or even overwhelmed by salient objects in the image. And the insufficient visual cues issue can not be handled by the cross-modal fusion mechanisms as done in previous work. In this paper, we tackle this problem from a novel perspective of enhancing the visual information for the referents by devising a Two-stage Visual cues enhancement Network (TV-Net), where a novel Retrieval and Enrichment Scheme (RES) and an Adaptive Multi-resolution feature Fusion (AMF) module are proposed. Specifically, RES retrieves the most relevant image from an external data pool with regard to both the visual and textual similarities, and then enriches the visual information of the referent with the retrieved image for better multimodal feature learning. AMF further enhances the visual detailed information by incorporating the highresolution feature maps from lower convolution layers of the image. Through the two-stage enhancement, our proposed TV-Net enjoys better performances in learning fine-grained matching behaviors between the natural language expression and image, especially when the visual information of the referent is inadequate, thus produces better segmentation results. Extensive experiments are conducted to validate the effectiveness of the proposed method on the RIS task, with our proposed TV-Net surpassing the state-of-theart approaches on four benchmark datasets. Our code is available at: https://github.com/SxJyJay/TV-Net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 被引用 48 次
- Temporal Collection and Distribution for Referring Video Object SegmentationJiajin Tang, Ge Zheng, Sibei YangICCV 2023 · 被引用 44 次
- Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal ModelsYang Jiao, Shaoxiang Chen, Zequn Jie, Jingjing Chen 等NeurIPS 2024 · 被引用 27 次
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang 等ACM MM 2022 · 被引用 12 次
- Suspected Objects Matter: Rethinking Model's Prediction for One-stage Visual GroundingYang Jiao, Zequn Jie, Jingjing Chen, Lin Ma 等ACM MM 2023 · 被引用 7 次
它引用的顶会 Paper6
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- See-Through-Text Grouping for Referring Image SegmentationDing-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen 等ICCV 2019 · 被引用 153 次
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu 等ACM MM 2020 · 被引用 74 次
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Bi-Directional Relationship Inferring Network for Referring Image SegmentationZhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang 等CVPR 2020
相关 Paper
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie 等ACM MM 2023 · 被引用 34 次
- Encoder Fusion Network With Co-Attention Embedding for Referring Image SegmentationGuang Feng, Zhiwei Hu, Lihe Zhang, Huchuan LuCVPR 2021
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang 等CVPR 2021
- Bottom-Up Shift and Reasoning for Referring Image SegmentationSibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou 等CVPR 2021
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu 等CVPR 2024
