Explicit Retrieval, Implicit Cognition: Towards Personalized Micro-video Popularity Prediction via Dual Latent Memory
Zhangtao Cheng, Bo Chen, Meihui Zhong, Ting Zhong, Bing Xia, Fan Zhou
Abstract
Micro-video popularity prediction (MVPP) is crucial for product marketing and recommendation. Existing works typically predict popularity by using independent modality encoders, but are limited by a personalized reasoning bottleneck: a tendency to lose grounding in multimodal evidence and a lack of personalized experience during reasoning. Inspired by human memory theory, we propose CLAMP, a cognition-memory-aligned framework that synergizes a large multimodal model (LMM) with dual latent memory, comprising a short-term module for content-centric perceptual retention and a long-term module for user-centric semantic consolidation. To achieve personalization, CLAMP develops a flexible memory retrieval mechanism governed by a familiarity-recollection process: familiarity enables fast recognition, while recollection supports deliberate, chain-like episodic retrieval. This mechanism adaptively switches between dual retrieval pathways to search for relevant memory knowledge. Furthermore, to invoke memory, CLAMP designs a dynamic invocation module that adaptively generates latent memory tokens to retain both perceptual evidence and user-level semantics. These memories are seamlessly invoked during inference, allowing the LMM to maintain perceptual fidelity and personalized semantics throughout both thinking and generation. Extensive experiments show that CLAMP consistently outperforms all baselines, validating its efficacy in personalized MVPP reasoning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- In-context Prompt-augmented Micro-video Popularity PredictionZhangtao Cheng, Jiao Li, Jian Lang, Ting Zhong et al.AAAI 2025 · 3 citations
- PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought ProcessXinliang Frederick Zhang, Nicholas Beauchamp, Lu WangEMNLP 2025
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen et al.CVPR 2026 · 30 citations
- Seeing the Unseen in Micro-Video Popularity Prediction: Self-Correlation Retrieval for Missing Modality GenerationZhangtao Cheng, Jian Lang, Ting Zhong, Fan ZhouKDD 2025 · 5 citations
- From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video AgentsNiu Lian, Yuting Wang, Hanshu Yao, Jinpeng Wang et al.ACL 2026 · 7 citations
