FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning
Wei Yang, Rui Zhong, Yiqun Chen, Shixuan Li, Heng Ping, Chi Lu, Peng Jiang
Abstract
Multimodal recommendation aims to enhance user preference modeling by leveraging rich item content such as images and text. Yet dominant systems fuse modalities in the spatial domain, obscuring the frequency structure of signals and amplifying misalignment and redundancy. We adopt a spectral information-theoretic view and show that, under an orthogonal transform that approximately block-diagonalizes bandwise covariances, the Gaussian Information Bottleneck objective decouples across frequency bands, providing a principled basis for separate-then-fuse paradigm. Building on this foundation, we propose FITMM, a Frequency-aware Information-Theoretic framework for multimodal recommendation. FITMM constructs graph-enhanced item representations, performs modality-wise spectral decomposition to obtain orthogonal bands, and forms lightweight within-band multimodal components. A residual, task-adaptive gate aggregates bands into the final representation. To control redundancy and improve generalization, we regularize training with a frequency-domain IB term that allocates capacity across bands (Wiener-like shrinkage with shut-off of weak bands). We further introduce a cross-modal spectral consistency loss that aligns modalities within each band. The model is jointly optimized with the standard recommendation loss. Extensive experiments on three real-world datasets demonstrate that FITMM consistently and significantly outperforms advanced baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e24bd0a-a799-4383-917b-08066787193dCited by top-tier papers1
Ask how each one uses itBuilds on25
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He et al.SIGIR 2021 · 1,476 citations
- Graph Neural Networks Exponentially Lose Expressive Power for Node ClassificationKenta Oono, Taiji SuzukiICLR 2020 · 864 citations
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen et al.SIGIR 2022 · 658 citations
Related papers
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- I3-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality RecommendationHuilin Chen, Miaomiao Cai, Fan Liu, Zhiyong Cheng et al.ACM MM 2025 · 1 citation
- Enhancing Multimodal Recommendation via Multimodal Representation Calibration in Spectral DomainMinghui Wang, Tingting Zhang, Yu Li, Yi ChangKDD 2026
- Frequency-refined Graph Convolution Network with Cross-modal Wavelet Denoising for RecommendationFeiyu Peng, Chaobo He, Junwei Cheng, Huijuan Hu et al.ACM MM 2025 · 6 citations
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva et al.ICML 2026
