Learning to Memorize with Attributive and Associative Memory for Online Test-Time Adaptation of Vision-Language Models
Yuchao Zhang, Hao Wang, Fan Zhang, QIRUI MI, Mengyue Yang, Yisen Wang, Jun Wang, Haoxuan Li, Zhouchen Lin
Abstract
Memory-based test-time adaptation (TTA) assigns streaming test samples into class-specific memory slots based on pseudo-labels predicted by models like CLIP, and retrieves them to facilitate subsequent predictions under distribution shift. However, this process introduces two challenges: (1) Each sample is hard-assigned to a single class based on CLIP's prediction, where inaccurate CLIP prediction leads to memory contamination that biases subsequent prediction. (2) Samples are evicted under biased selection due to fixed memory capacity, which risks discarding informative samples and undermining the efficacy of the memory. To address these challenges, we propose A²Memory (Attributive-Associative Memory for Test-time Adaptation). For challenge (1), we propose Attribute-centric Memory Construction that builds prior textual representations from class-shared representative and diverse visual attributes, and applies soft assignment to generate surrogate visual representations. For challenge (2), we design Class-wise Associative Memory that dynamically compresses streaming samples into fixed-capacity memory through gradient-based optimization and data-dependent retention, then retrieves sample-adaptive class prototypes for reliable inference. Extensive experiments demonstrate consistent improvements over state-of-the-art methods across 15 benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2dc9f9d-2c06-40fd-bcca-295ad986b8c4Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
Related papers
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language ModelsSi Chen, Yujia Chen, Xiaotian Yin, Xin Liu et al.ACM MM 2025 · 1 citation
- Bayesian Test-Time Adaptation for Vision-Language ModelsLihua Zhou, Mao Ye, Shuaifeng Li, Nianxin Li et al.CVPR 2025
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
- BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingTaolin Zhang, Jinpeng Wang, Hang Guo, Tao Dai et al.NeurIPS 2024 · 30 citations
