Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation
Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang, Hui Xiong
Abstract
Multi-modal sequential recommendation (SR) leverages multi-modal data to learn more comprehensive item features and user preferences than traditional SR methods, which has become a critical topic in both academia and industry. Existing methods typically focus on enhancing multi-modal information utility through adaptive modality fusion to capture the evolving of user preference from user-item interaction sequences. However, most of them overlook the interference caused by redundant interest-irrelevant information contained in rich multi-modal data. Additionally, they primarily rely on implicit temporal information based solely on chronological ordering, neglecting explicit temporal signals that could more effectively represent dynamic user interest over time. To address these limitations, we propose a Hierarchical time-aware Mixture of experts for multi-modal Sequential Recommendation (HM4SR) with a two-level Mixture of Experts (MoE) and a multi-task learning strategy. Specifically, the first MoE, named Interactive MoE, extracts essential user interest-related information from the multi-modal data of each item. Then, the second MoE, termed Temporal MoE, captures user dynamic interests by introducing explicit temporal embeddings from timestamps in modality encoding. To further address data sparsity, we propose three auxiliary supervision tasks: sequence-level category prediction (CP) for item feature understanding, contrastive learning on ID (IDCL) to align sequence context with user interests, and placeholder contrastive learning (PCL) to integrate temporal information with modalities for dynamic interest modeling. Extensive experiments on four public datasets verify the effectiveness of HM4SR compared to several state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b5d3598-8991-45af-9c9d-be2771bfce39Cited by top-tier papers10
- Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence RecommendationXiao Lin, Zhicheng Tang, Weilin Cong, Mengyue Hang et al.WWW 2026 · 3 citations
- FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential RecommendationMaolin Wang, Yutian Xiao, Binhao Wang, Sheng Zhang et al.KDD 2025 · 3 citations
- CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ICDE 2026 · 1 citation
- Personalized Parameter-Efficient Fine-Tuning of Foundation Models for Multimodal RecommendationSunwoo Kim, Hyunjin Hwang, Kijung ShinWWW 2026 · 1 citation
- Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-ExpertsXing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang et al.ICLR 2026 · 1 citation
Builds on26
- Filter-enhanced MLP is All You Need for Sequential RecommendationKun Zhou, Hui Yu, Wayne Xin Zhao, Ji-Rong WenWWW 2022 · 411 citations
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li et al.KDD 2022 · 245 citations
- Noninvasive Self-attention for Side Information Fusion in Sequential RecommendationChang Liu, Xiaoguang Li, Guohao Cai, Zhenhua Dong et al.AAAI 2021 · 177 citations
- Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge GraphsLiyi Chen, Panrong Tong, Zhongming Jin, Ying Sun et al.NeurIPS 2024 · 160 citations
- GeomGCL: Geometric Graph Contrastive Learning for Molecular Property PredictionShuangli Li, Jingbo Zhou, Tong Xu, Dejing Dou et al.AAAI 2022 · 158 citations
Related papers
- Capturing Dynamic User Interests Under Modality Imbalance for Multimodal Sequential RecommendationZilong Li, Jia Zhu, Chenglei Huang, Zhangze Chen et al.AAAI 2026
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang et al.ACM MM 2023 · 62 citations
- M²VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item RecommendationChuan He, Yongchao Liu, Qiang Li, Chuntao Hong et al.AAAI 2026 · 1 citation
- CMCLRec: Cross-modal Contrastive Learning for User Cold-start Sequential RecommendationXiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang et al.SIGIR 2024 · 56 citations
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
