Explicit Retrieval, Implicit Cognition: Towards Personalized Micro-video Popularity Prediction via Dual Latent Memory
Zhangtao Cheng, Bo Chen, Meihui Zhong, Ting Zhong, Bing Xia, Fan Zhou
摘要
Micro-video popularity prediction (MVPP) is crucial for product marketing and recommendation. Existing works typically predict popularity by using independent modality encoders, but are limited by a personalized reasoning bottleneck: a tendency to lose grounding in multimodal evidence and a lack of personalized experience during reasoning. Inspired by human memory theory, we propose CLAMP, a cognition-memory-aligned framework that synergizes a large multimodal model (LMM) with dual latent memory, comprising a short-term module for content-centric perceptual retention and a long-term module for user-centric semantic consolidation. To achieve personalization, CLAMP develops a flexible memory retrieval mechanism governed by a familiarity-recollection process: familiarity enables fast recognition, while recollection supports deliberate, chain-like episodic retrieval. This mechanism adaptively switches between dual retrieval pathways to search for relevant memory knowledge. Furthermore, to invoke memory, CLAMP designs a dynamic invocation module that adaptively generates latent memory tokens to retain both perceptual evidence and user-level semantics. These memories are seamlessly invoked during inference, allowing the LMM to maintain perceptual fidelity and personalized semantics throughout both thinking and generation. Extensive experiments show that CLAMP consistently outperforms all baselines, validating its efficacy in personalized MVPP reasoning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- In-context Prompt-augmented Micro-video Popularity PredictionZhangtao Cheng, Jiao Li, Jian Lang, Ting Zhong 等AAAI 2025 · 被引用 3 次
- PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought ProcessXinliang Frederick Zhang, Nicholas Beauchamp, Lu WangEMNLP 2025
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen 等CVPR 2026 · 被引用 30 次
- Seeing the Unseen in Micro-Video Popularity Prediction: Self-Correlation Retrieval for Missing Modality GenerationZhangtao Cheng, Jian Lang, Ting Zhong, Fan ZhouKDD 2025 · 被引用 5 次
- From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video AgentsNiu Lian, Yuting Wang, Hanshu Yao, Jinpeng Wang 等ACL 2026 · 被引用 7 次
