Multi-Task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, Rongrong Ji
Abstract
Referring expression comprehension (REC) and segmentation (RES) are two highly-related tasks, which both aim at identifying the referent according to a natural language expression. In this paper, we propose a novel Multi-task Collaborative Network (MCN) 1 to achieve a joint learning of REC and RES for the first time. In MCN, RES can help REC to achieve better language-vision alignment, while REC can help RES to better locate the referent. In addition, we address a key challenge in this multi-task setup, i.e., the prediction conflict, with two innovative designs namely, Consistency Energy Maximization (CEM) and Adaptive Soft Non-Located Suppression (ASNLS). Specifically, CEM enables REC and RES to focus on similar visual regions by maximizing the consistency energy between two tasks. ASNLS supresses the response of unrelated regions in RES based on the prediction of REC. To validate our model, we conduct extensive experiments on three benchmark datasets of REC and RES, i.e., RefCOCO, RefCOCO+ and Ref-COCOg. The experimental results report the significant performance gains of MCN over all existing methods, i.e., up to +7.13% for REC and +11.50% for RES over SOTA, which well confirm the validity of our model for joint REC and RES learning. * Equal Contribution. † Corresponding Author. 1 Source codes and pretrained backbone are available at : https:// github.com/luogen1996/MCN "a half horse." (a) Illustration of Referring Expression Comprehension (REC) and Segmentation (RES). Referring Expression Segmentation Referring Expression Comprehension "person on scooter wearing black helmet and has black backpack" "the cat right in front of the window.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31301d0b-3e6a-40c5-bdbd-86bb2e5576a3Cited by top-tier papers125
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu et al.ICCV 2023 · 327 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
Builds on4
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- Zero-Shot Grounding of Objects From Natural Language QueriesArka Sadhu, Kan Chen, Ram NevatiaICCV 2019 · 176 citations
Related papers
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationSilin Cheng, Yang Liu, Xinwei He, Sébastien Ourselin et al.CVPR 2025
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 10 citations
- Cascade Grouped Attention Network for Referring Expression SegmentationGen Luo, Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun et al.ACM MM 2020 · 142 citations
- Whether you can locate or not? Interactive Referring Expression GenerationFulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie WangACM MM 2023 · 6 citations
- Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment NetworkZiling Huang, Shin'ichi SatohEMNLP 2023 · 4 citations
