VBF++: Variational Bayesian Fusion with Context-Aware Priors and Recommendation-Guided Adversarial Refinement for Multimodal Video Recommendation
Ziyi Cao, Rui Liu, Yong Chen
摘要
Multimodal video recommendation systems face fundamental challenges in determining optimal fusion strategies across diverse content types and user preferences. Existing methods suffer from two critical limitations: (1) their fusion strategies are guided by context-agnostic priors that ignore the semantic structure of content, assuming the same simple distribution (typically a standard multivariate Gaussian prior) governs optimal fusion for all video types, and (2) their optimization objectives, particularly the Evidence Lower Bound (ELBO), are misaligned with the final recommendation goal, optimizing for feature reconstruction rather than ranking performance. To address these fundamental issues, this work proposes VBF++, a novel framework that introduces context-aware structured priors and recommendation-guided adversarial refinement. First, the method designs context-aware priors that learn cluster-specific distributions based on video semantic categories, replacing uninformative priors with structured, content-aware prior distributions. Second, it introduces a Recommendation-Guided Adversarial Refinement (RAR) paradigm that explicitly steers the learning process towards generating recommendation-optimal fusion strategies, resolving the objective misalignment inherent in variational learning. Enhanced with domain-adaptive meta-learning, extensive experiments on three real-world datasets demonstrate consistent improvements of 4.7-8.3 percent in Precision@10 over state-of-the-art methods. Analysis reveals that learned fusion strategies exhibit semantically meaningful patterns, prioritizing visual features for action content, acoustic information for music videos, and textual descriptions for documentary material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Learning Intents behind Interactions with Knowledge Graph for RecommendationXiang Wang, Tinglin Huang, Dingxian Wang, Yancheng Yuan 等WWW 2021 · 被引用 584 次
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He 等ACM MM 2020 · 被引用 374 次
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 被引用 256 次
- Modality-Independent Graph Neural Networks with Global Transformers for Multimodal RecommendationJun Hu, Bryan Hooi, Bingsheng He, Yinwei WeiAAAI 2025 · 被引用 31 次
相关 Paper
- Multi-modal Bipartite Graph Structure Learning with Information Bottleneck for Micro-video RecommendationYing He, Desheng Cai, Shengsheng Qian, Quan Fang 等WWW 2026
- Capturing Dynamic User Interests Under Modality Imbalance for Multimodal Sequential RecommendationZilong Li, Jia Zhu, Chenglei Huang, Zhangze Chen 等AAAI 2026
- In-context Prompt-augmented Micro-video Popularity PredictionZhangtao Cheng, Jiao Li, Jian Lang, Ting Zhong 等AAAI 2025 · 被引用 3 次
- Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingZecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq 等SIGIR 2025 · 被引用 6 次
- Parameter-Efficient Variational AutoEncoder for Multimodal Multi-Interest RecommendationNhu-Thuat Tran, Hady W. LauwACM MM 2025
