Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
Yangyang Liu, Yuhao Wang, Pingping Zhang
Abstract
Multi-modal object Re-IDentification (ReID) is devoted to retrieving specific objects through the exploitation of complementary multi-modal image information. Existing methods mainly concentrate on the fusion of multi-modal features, yet neglecting the background interference. Besides, current multi-modal fusion methods often focus on aligning modality pairs but suffer from multi-modal consistency alignment. To address these issues, we propose a novel selective interaction and global-local alignment framework called Signal for multi-modal object ReID. Specifically, we first propose a Selective Interaction Module (SIM) to select important patch tokens with intra-modal and inter-modal information. These important patch tokens engage in the interaction with class tokens, thereby yielding more discriminative features. Then, we propose a Global Alignment Module (GAM) to simultaneously align multi-modal features by minimizing the volume of 3D polyhedra in the gramian space. Meanwhile, we propose a Local Alignment Module (LAM) to align local features in a shift-aware manner. With these modules, our proposed framework could extract more discriminative features for object ReID. Extensive experiments on three multi-modal object ReID benchmarks (i.e., RGBNT201, RGBNT100, MSVR310) validate the effectiveness of our method. The source code is available at https://github.com/010129/Signal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7a64a27-8875-4bce-b57c-ebddc905324cBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- HAT: Hierarchical Aggregation Transformers for Person Re-identificationGuowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan LuACM MM 2021 · 159 citations
Related papers
- Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu et al.CVPR 2024 · 42 citations
- STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-IdentificationXingguo Xu, Zhanyu Liu, Weixiang Zhou, Yuansheng Gao et al.AAAI 2026
- Multi-Modal Object Re-identification via Sparse Mixture-of-ExpertsYingying Feng, Jie Li, Chi Xie, Lei Tan et al.ICML 2025
- Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identificationZi Wang, Chenglong Li, Aihua Zheng, Ran He et al.AAAI 2022 · 61 citations
- FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-IdentificationXuanhao Qi, Tom Luan, Yukang Zhang, Jinkai Zheng et al.ICML 2026
