Spatially-Regularized Entropy for Discriminative Token Merging in Fine-Grained Re-Identification
Shangze Li, Yifan Xu, Jingmiao Liang, Yongfei Zhang, Yuzhuo Ma, Yingbo Qu
摘要
While Vision Transformers (ViTs) offer strong global modeling, their quadratic computational cost limits utility in latency-sensitive applications like person re-identification (ReID). Existing compression strategies, such as token pruning or generic merging, typically rely on coarse-grained criteria tailored for image classification. In fine-grained retrieval, these approaches often discard or smooth out subtle but discriminative local details. To resolve this, we propose SRE-Merge, a training-free framework designed for discriminative token compression. SRE-Merge injects spatial priors into the merging process through three mechanisms: (i) Spatial-Entropy Saliency Assessment (SES-Assess), which quantifies token importance as Spatial-Entropic Mass (SE-Mass) by coupling spatial structure with local attention entropy; (ii) Hybrid Context-Affinity Matching (HCA-Match), which guides precise pair selection by combining feature similarity with mass-derived context; and (iii) Energy-Preserving Weighted Fusion (EPW-Fuse), which incorporates SE-Mass weighting to counteract feature variance reduction. Extensive experiments on standard benchmarks show that SRE-Merge reduces GFLOPs of the base ViT model by about 24% while retaining competitive retrieval accuracy, establishing a superior accuracy-efficiency trade-off.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- Pose-Guided Feature Alignment for Occluded Person Re-IdentificationJiaxu Miao, Yu Wu, Ping Liu, Yuhang Ding 等ICCV 2019 · 被引用 589 次
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu 等CVPR 2022 · 被引用 251 次
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo 等AAAI 2022 · 被引用 248 次
- Feature Erasing and Diffusion Network for Occluded Person Re-IdentificationZhikang Wang, Feng Zhu, Shixiang Tang, Rui Zhao 等CVPR 2022 · 被引用 185 次
相关 Paper
- Saliency-Driven Token Merging for Vision TransformersWeiying Xie, Xiaoyu Chen, Xin Zhang, Chenhe Hao 等CVPR 2026
- TF-ATM: Training-Free Adaptive Token MergingXin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen 等ACM MM 2025
- vid-TLDR: Training Free Token merging for Light-Weight Video TransformerJoonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi 等CVPR 2024
- SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-IdentificationHuiyuan Huang, SANG MIN YOONCVPR 2026
- Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersSifan Long, Zhen Zhao, Jimin Pi, Shengsheng Wang 等CVPR 2023
