Spatially-Regularized Entropy for Discriminative Token Merging in Fine-Grained Re-Identification
Shangze Li, Yifan Xu, Jingmiao Liang, Yongfei Zhang, Yuzhuo Ma, Yingbo Qu
Abstract
While Vision Transformers (ViTs) offer strong global modeling, their quadratic computational cost limits utility in latency-sensitive applications like person re-identification (ReID). Existing compression strategies, such as token pruning or generic merging, typically rely on coarse-grained criteria tailored for image classification. In fine-grained retrieval, these approaches often discard or smooth out subtle but discriminative local details. To resolve this, we propose SRE-Merge, a training-free framework designed for discriminative token compression. SRE-Merge injects spatial priors into the merging process through three mechanisms: (i) Spatial-Entropy Saliency Assessment (SES-Assess), which quantifies token importance as Spatial-Entropic Mass (SE-Mass) by coupling spatial structure with local attention entropy; (ii) Hybrid Context-Affinity Matching (HCA-Match), which guides precise pair selection by combining feature similarity with mass-derived context; and (iii) Energy-Preserving Weighted Fusion (EPW-Fuse), which incorporates SE-Mass weighting to counteract feature variance reduction. Extensive experiments on standard benchmarks show that SRE-Merge reduces GFLOPs of the base ViT model by about 24% while retaining competitive retrieval accuracy, establishing a superior accuracy-efficiency trade-off.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 360b613b-1a33-4307-a86e-8f7ec9e18633Builds on16
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Pose-Guided Feature Alignment for Occluded Person Re-IdentificationJiaxu Miao, Yu Wu, Ping Liu, Yuhang Ding et al.ICCV 2019 · 589 citations
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo et al.AAAI 2022 · 248 citations
- Feature Erasing and Diffusion Network for Occluded Person Re-IdentificationZhikang Wang, Feng Zhu, Shixiang Tang, Rui Zhao et al.CVPR 2022 · 185 citations
Related papers
- Saliency-Driven Token Merging for Vision TransformersWeiying Xie, Xiaoyu Chen, Xin Zhang, Chenhe Hao et al.CVPR 2026
- TF-ATM: Training-Free Adaptive Token MergingXin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen et al.ACM MM 2025
- vid-TLDR: Training Free Token merging for Light-Weight Video TransformerJoonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi et al.CVPR 2024
- SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-IdentificationHuiyuan Huang, SANG MIN YOONCVPR 2026
- Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersSifan Long, Zhen Zhao, Jimin Pi, Shengsheng Wang et al.CVPR 2023
