Chain-of-Thought Guided Multi-Modal Object Re-Identification
Ya Gao, Shihao Li, Zhaojun Liu, Aihua Zheng, Chenglong Li, Jin Tang
Abstract
With the rise of visual-language models, multi-modal ReID retrieves specific targets by integrating different spectra and textual descriptions. Existing methods merely adopt descriptive representation learning for image-text, ignoring the relationships among the intrinsic logical hierarchies of semantic features. Since Chain-of-Thought (CoT) can provide textual logical context and enhance semantic perception in large-model reasoning, we propose CoT-ReID, a CoT-guided framework that injects the Multi-modal Large Language Models (MLLMs) reasoning into multi-modal ReID. Specifically, we simulate human-like joint visiontext logical decision-making, leveraging CoT textual logical reasoning to guide visual feature learning at the early, late and decision-making levels: we first embed the semantic reversion of CoT hierarchical reasoning into visual features to calibrate bottom-level features and highlight visual hierarchical reasoning, then take CoT hierarchical reasoning text as an anchor condition to constrain the consistency of visual cross-modal semantics, and finally embed logically reasoned text attribute features into multi-modal decision-making via CoT's hierarchical reasoning to provide logical support for selecting discriminative identity features. By constructing CoT textual benchmarks and our proposed modules, our framework generates more robust multi-modal features in complex scenarios, and comprehensive experiments on four datasets (RGBNT100, MSVR310, WMVeID863, RGBNT201) demonstrate the superiority of our method over state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69bab88d-606d-42d5-a9f0-e6b68f741ecaBuilds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
Related papers
- MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image RetrievalXuri Ge, Chunhao Wang, Xindi Wang, Zheyun Qin et al.WWW 2026
- Cantor: Inspiring Multimodal Chain-of-Thought of MLLMTimin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu et al.ACM MM 2024 · 20 citations
- Rationale-Enhanced Decoding for Multi-modal Chain-of-ThoughtShin'ya Yamaguchi, Kosuke Nishida, Daiki ChijiwaCVPR 2026
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li et al.ICLR 2026 · 55 citations
