ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identification
Aihua Zheng, Zhaojun Liu, Xixi Wan, Chenglong Li, Jin Tang, Yan Yan
摘要
Multi-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the difficulty of effectively leveraging reliable target-domain data to improve generalization. To address these challenges, we propose Proxy-driven Test-Time Training (ProxyTTT), a unified framework that enhances both multi-modal identity representation learning and model generalization. During training, we propose a Multi-Proxy Learning (MPL) mechanism to address the representation bias across different views and modalities. MPL disentangles fine-grained modality-specific and modality-common identity proxies as semantic anchors to align identity features across diverse perspectives and sensing modalities. This alignment strategy enables the model to learn robust and discriminative global identity representations under heterogeneous modality conditions. At test time, to reliably exploit target domain data, we propose Proxy-guided Entropy-based Selective Adaptation (PESA) for test-time training. Specifically, PESA leverages the semantic structure encoded by identity proxies to estimate prediction uncertainty via entropy, and selectively adapts the model using only high-confidence samples. This selective adaptation effectively mitigates the domain shift between training and deployment environments, improving the model’s generalization in real-world scenarios. Extensive experiments on four public multi-modal ReID benchmarks (RGBNT201, RGBNT100, MSVR310, and WMVeID863) demonstrate the effectiveness of ProxyTTT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
相关 Paper
- Heterogeneous Test-Time Training for Multi-Modal Person Re-identificationZi Wang, Huaibo Huang, Aihua Zheng, Ran HeAAAI 2024 · 被引用 22 次
- Object-Generalized Re-Identification: A Step Towards Universal Instance PerceptionShuoyi Chen, Yurui Wu, Mang YeCVPR 2026 · 被引用 1 次
- Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-IdentificationZhongao Zhou, Bin Yang, Wenke Huang, Jun Chen 等NeurIPS 2025 · 被引用 2 次
- Towards Modality-Agnostic Person Re-identification with Descriptive QueryCuiqun Chen, Mang Ye, Ding JiangCVPR 2023
- Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time AdaptationJiacheng Li, Songhe FengAAAI 2026 · 被引用 2 次
