FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification
Xuanhao Qi, Tom Luan, Yukang Zhang, Jinkai Zheng, su zhou, Shuwei Li, Lei Tan
Abstract
Despite significant progress in multi-modal Re-Identification (ReID), existing methods tend to emphasize low-frequency cues. Consequently, they focus on attributes such as color, illumination, and coarse appearance, while overlooking mid- and high-frequency structures that encode geometric, textural, and identity-discriminative details. This imbalance leads to incomplete spectral representations and unstable cross-modal alignment. To overcome these limitations, we introduce FUSE, a frequency-domain framework that reformulates multi-modal ReID as a two-stage process of spectral disentanglement and energy alignment. The proposed Spectral Decomposition Module (SDM) adaptively partitions features into low, mid, and high-frequency subspaces, enabling hierarchical spectral modeling. The Cross-Modal Alignment Module (CAM) further enforces energy alignment and subspace complementarity across modalities via frequency-consistency regularization. In addition, FUSE incorporates learnable frequency modulation to enhance robustness under varying illumination and heterogeneous sensor conditions. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 show that FUSE achieves 9.1% mAP and 9.5% Rank-1 improvements, establishing an interpretable frequency-domain paradigm for multi-modal representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Decoupled Contrastive Multi-View Clustering with High-Order Random WalksYiding Lu, Yijie Lin, Mouxing Yang, Dezhong Peng et al.AAAI 2024 · 107 citations
- Multi-Spectral Vehicle Re-Identification: A ChallengeHongchao Li, Chenglong Li, Xianpeng Zhu, Aihua Zheng et al.AAAI 2020 · 81 citations
Related papers
- Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-IdentificationJIan Yu, Yujian Feng, Shuai You, Zhongkai Zhou et al.CVPR 2026
- MFEN: Multi-Frequency Expert Network for Visible-Infrared Person Re-IDXulin Li, Yan Lu, Bin Liu, Qinhong Yang et al.CVPR 2026 · 2 citations
- Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-IdentificationYangyang Liu, Yuhao Wang, Pingping ZhangAAAI 2026
- FAMRD: Frequency-Aware Multimodal Reverse Distillation for Industrial Anomaly DetectionQiyin Zhong, Xianglin Qiu, Xiaolei Wang, Zhen Zhang et al.ACM MM 2025
- Robust Multi-Modality Person Re-identificationAihua Zheng, Zi Wang, Zi-Han Chen, Chenglong Li et al.AAAI 2021 · 79 citations
