Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation
Wei Yang, Rui Zhong, Yiqun Chen, Chi Lu, Peng Jiang
Abstract
Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over user-item graphs. These issues are often exacerbated by naive fusion or shallow modeling strategies, leading to degraded generalization and poor robustness. While recent work has explored the frequency domain as a lens to separate stable from noisy signals, most methods rely on static filtering or reweighting, lacking the ability to reason over spectral structure or adapt to modality-specific reliability. To address these challenges, we propose a Structured Spectral Reasoning (SSR) framework for frequency-aware multimodal recommendation. Our method follows a four-stage pipeline: (i) Decompose graph-based multimodal signals into spectral bands via graph-guided transformations to isolate semantic granularity; (ii) Modulate band-level reliability with spectral band masking, a training-time masking with representation-consistency objective that suppresses brittle frequency components; (iii) Fuse complementary frequency cues using hyperspectral reasoning with low-rank cross-band interaction; and (iv) Align modality-specific spectral features via contrastive regularization to promote semantic and structural consistency. Experiments on three real-world benchmarks show consistent gains over strong baselines, particularly under sparse and cold-start settings. Additional analyses indicate that structured spectral modeling improves robustness and provides clearer diagnostics of how different bands contribute to performance. The code is available at https://github.com/llm-ml/SSR.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8da3f1d7-aed9-4fe7-80f5-bd13caa2ec36Cited by top-tier papers2
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li et al.ACM MM 2025 · 6 citations
- TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal RecommendationWei Yang, Rui Zhong, Zihan Lin, Xiaodan Wang et al.SIGIR 2026
Builds on36
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen et al.SIGIR 2022 · 658 citations
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He et al.SIGIR 2020 · 621 citations
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu et al.ACM MM 2021 · 350 citations
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
Related papers
- Frequency-refined Graph Convolution Network with Cross-modal Wavelet Denoising for RecommendationFeiyu Peng, Chaobo He, Junwei Cheng, Huijuan Hu et al.ACM MM 2025 · 6 citations
- DIGEST: Dynamic Graph Refinement with Dual Contrastive Semantic Transfer for Multimodal RecommendationXiangyu Sai, Meysam Madadi, Sergio Escalera, Yong XuSIGIR 2026
- Dynamic Spectral Denoising with Global-Context Attention for Multi-Behavior RecommendationMiaomiao Cai, Yunshan Ma, Fangqi Zhu, Junfeng Fang et al.KDD 2026
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
- SGP4SR: Separated-Modality Guided User Preference Learning for Multimodal Sequential RecommendationChanghong Li, Zhiqiang Guo, Guohui Li, Zhong Yang et al.AAAI 2026
