Language-Guided Salient Object Ranking
Fang Liu, Yuhao Liu, Ke Xu, Shuquan Ye, Gerhard Petrus Hancke, Rynson W. H. Lau
Abstract
Salient Object Ranking (SOR) aims to study human attention shifts across different objects in the scene. It is a challenging task, as it requires comprehension of the relations among the salient objects in the scene. However, existing works often overlook such relations or model them implicitly. In this work, we observe that when Large Vision-Language Models (LVLMs) describe a scene, they usually focus on the most salient object first, and then discuss the relations as they move on to the next (less salient) one. Based on this observation, we propose a novel Language-Guided Salient Object Ranking approach (named LG-SOR), which utilizes the internal knowledge within the LVLM-generated language descriptions, i.e., semantic relation cues and the implicit entity order cues, to facilitate saliency ranking. Specifically, we first propose a novel Text-Guided Visual Modulation (TGVM) module to incorporate semantic information in the description for saliency ranking. TGVM controls the flow of linguistic information to the visual features, suppresses noisy background image features, and enables the propagation of useful textual features. We then propose a novel Text-Aware Visual Reasoning (TAVR) module to enhance model reasoning in object ranking, by explicitly learning a multimodal graph based on the entity and relation cues derived from the description. Extensive experiments demonstrate superior performances of our model on two SOR benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa1cab8c-3306-4441-bc09-1711bce37dabCited by top-tier papers4
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingZaiquan Yang, Yuhao Liu, Gerhard P. Hancke, Rynson W. H. LauNeurIPS 2025 · 10 citations
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionJi Du, Xin Wang, Fangwei Hao, Mingyang Yu et al.ICCV 2025 · 2 citations
- Salient Object Ranking via Cyclical Perception-Viewing Interaction ModelingRongjin Guo, Ke Xu, Rynson W. H. LauICLR 2026
- Probabilistic Salient Object RankingRongjin Guo, Guan Huankang, Rynson W LauICML 2026
Builds on37
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with DiffusionYunlong Tang, Gen Zhan, Li Yang, Yiting Liao et al.AAAI 2025 · 16 citations
- Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context LearningYanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li et al.AAAI 2026 · 9 citations
- Bi-directional Object-Context Prioritization Learning for Saliency RankingXin Tian, Ke Xu, Xin Yang, Lin Du et al.CVPR 2022 · 33 citations
- Chain-of-Thought Guided Multi-Modal Object Re-IdentificationYa Gao, Shihao Li, Zhaojun Liu, Aihua Zheng et al.CVPR 2026
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
