CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
Yunlong Tang, Gen Zhan, Li Yang, Yiting Liao, Chenliang Xu
摘要
Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like memory and cognition. Among these top-down influences, language plays a crucial role in guiding attention by shaping how visual information is interpreted. Existing methods primarily focus on modeling perceptual information while neglecting the reasoning process facilitated by language, where ranking cues are crucial outcomes of this process and practical guidance for saliency prediction. In this paper, we propose CaRDiff (Caption, Rank, and generate with Diffusion), a framework that imitates the process by integrating multimodal large language model (MLLM), a grounding module, and a diffusion model, to enhance video saliency prediction. Specifically, we introduce a novel prompting method VSOR-CoT (Video Slient Object Ranking Chain of Thought), which utilizes an MLLM with a grounding module to caption video content and infer salient objects along with their rankings and positions. This process derives ranking maps that can be sufficiently leveraged by the diffusion model to accurately decode the saliency maps for the given video. Extensive experiments showcase the effectiveness of VSOR-CoT in improving the performance of video saliency prediction. The proposed CaRDiff performs better than state-of-the-art models on the MVS dataset and demonstrates cross-dataset capabilities on the DHF1k dataset through zero-shot evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language ModelsZhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu 等EMNLP 2025 · 被引用 5 次
- Cut to the Chase: Training-free Multimodal Summarization via Chain-of-EventsXiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang 等CVPR 2026 · 被引用 4 次
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
- Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentYizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim 等ICLR 2025
- Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture RepresentationPinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Garrido 等ACM MM 2025
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Language-Guided Salient Object RankingFang Liu, Yuhao Liu, Ke Xu, Shuquan Ye 等CVPR 2025
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li 等CVPR 2026 · 被引用 5 次
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Salient Object Ranking via Cyclical Perception-Viewing Interaction ModelingRongjin Guo, Ke Xu, Rynson W. H. LauICLR 2026
