VBF++: Variational Bayesian Fusion with Context-Aware Priors and Recommendation-Guided Adversarial Refinement for Multimodal Video Recommendation
Ziyi Cao, Rui Liu, Yong Chen
Abstract
Multimodal video recommendation systems face fundamental challenges in determining optimal fusion strategies across diverse content types and user preferences. Existing methods suffer from two critical limitations: (1) their fusion strategies are guided by context-agnostic priors that ignore the semantic structure of content, assuming the same simple distribution (typically a standard multivariate Gaussian prior) governs optimal fusion for all video types, and (2) their optimization objectives, particularly the Evidence Lower Bound (ELBO), are misaligned with the final recommendation goal, optimizing for feature reconstruction rather than ranking performance. To address these fundamental issues, this work proposes VBF++, a novel framework that introduces context-aware structured priors and recommendation-guided adversarial refinement. First, the method designs context-aware priors that learn cluster-specific distributions based on video semantic categories, replacing uninformative priors with structured, content-aware prior distributions. Second, it introduces a Recommendation-Guided Adversarial Refinement (RAR) paradigm that explicitly steers the learning process towards generating recommendation-optimal fusion strategies, resolving the objective misalignment inherent in variational learning. Enhanced with domain-adaptive meta-learning, extensive experiments on three real-world datasets demonstrate consistent improvements of 4.7-8.3 percent in Precision@10 over state-of-the-art methods. Analysis reveals that learned fusion strategies exhibit semantically meaningful patterns, prioritizing visual features for action content, acoustic information for music videos, and textual descriptions for documentary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d7a1bec-281e-45a3-8e6c-07ab0641ea1cBuilds on6
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Learning Intents behind Interactions with Knowledge Graph for RecommendationXiang Wang, Tinglin Huang, Dingxian Wang, Yancheng Yuan et al.WWW 2021 · 584 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- Modality-Independent Graph Neural Networks with Global Transformers for Multimodal RecommendationJun Hu, Bryan Hooi, Bingsheng He, Yinwei WeiAAAI 2025 · 31 citations
Related papers
- Multi-modal Bipartite Graph Structure Learning with Information Bottleneck for Micro-video RecommendationYing He, Desheng Cai, Shengsheng Qian, Quan Fang et al.WWW 2026
- Capturing Dynamic User Interests Under Modality Imbalance for Multimodal Sequential RecommendationZilong Li, Jia Zhu, Chenglei Huang, Zhangze Chen et al.AAAI 2026
- In-context Prompt-augmented Micro-video Popularity PredictionZhangtao Cheng, Jiao Li, Jian Lang, Ting Zhong et al.AAAI 2025 · 3 citations
- Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingZecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq et al.SIGIR 2025 · 6 citations
- Parameter-Efficient Variational AutoEncoder for Multimodal Multi-Interest RecommendationNhu-Thuat Tran, Hady W. LauwACM MM 2025
