Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
Pingping Zhang, Yuhao Wang, Yang Liu, Zhengzheng Tu, Huchuan Lu
Abstract
Single-modal object re-identification (ReID) faces great challenges in maintaining robustness within complex visual scenarios. In contrast, multi-modal object ReID utilizes complementary information from diverse modalities, showing great potentials for practical applications. How-ever, previous methods may be easily affected by irrele-vant backgrounds and usually ignore the modality gaps. To address above issues, we propose a novel learning frame-work named EDITOR to select diverse tokens from vision Transformers for multi-modal object ReID. We be-gin with a shared vision Transformer to extract tokenized features from different input modalities. Then, we intro-duce a Spatial-Frequency Token Selection (SFTS) module to adaptively select object-centric tokens with both spa-tial and frequency information. Afterwards, we employ a Hierarchical Masked Aggregation (HMA) module to fa-cilitate feature interactions within and across modalities. Finally, to further reduce the effect of backgrounds, we propose a Background Consistency Constraint (BCC) and an Object-Centric Feature Refinement (OCFR). They are formulated as two new loss functions, which improve the feature discrimination with background suppression. As a result, our framework can generate more discriminative features for multi-modal object ReID. Extensive ex-periments on three multi-modal ReID benchmarks verify the effectiveness of our methods. The code is available at https://github.com/924973292/EDITOR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 951e9001-e2a1-4dd7-a691-f98a5fdc8356Cited by top-tier papers26
- Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationJiangming Shi, Xiangbo Yin, Yachao Zhang, Zhizhong Zhang et al.NeurIPS 2024 · 36 citations
- DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-IdentificationYuhao Wang, Yang Liu, Aihua Zheng, Pingping ZhangAAAI 2025 · 31 citations
- MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptYuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu et al.AAAI 2025 · 30 citations
- Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationXiangbo Yin, Jiangming Shi, Yachao Zhang, Yang Lu et al.ACM MM 2024 · 28 citations
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
Related papers
- TOP-ReID: Multi-Spectral Object Re-identification with Token PermutationYuhao Wang, Xuehu Liu, Pingping Zhang, Hu Lu et al.AAAI 2024 · 49 citations
- Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-IdentificationYangyang Liu, Yuhao Wang, Pingping ZhangAAAI 2026
- STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-IdentificationXingguo Xu, Zhanyu Liu, Weixiang Zhou, Yuansheng Gao et al.AAAI 2026
- TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-IdentificationXiao Wang, Lekai Liu, Bin Yang, Mang Ye et al.AAAI 2025 · 8 citations
- Miss-ReID: Delivering Robust Multi-Modality Object Re-Identification Despite Missing ModalitiesRuida XiNeurIPS 2025 · 4 citations
