Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation
Wei Yang, Rui Zhong, Yiqun Chen, Chi Lu, Peng Jiang
摘要
Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over user-item graphs. These issues are often exacerbated by naive fusion or shallow modeling strategies, leading to degraded generalization and poor robustness. While recent work has explored the frequency domain as a lens to separate stable from noisy signals, most methods rely on static filtering or reweighting, lacking the ability to reason over spectral structure or adapt to modality-specific reliability. To address these challenges, we propose a Structured Spectral Reasoning (SSR) framework for frequency-aware multimodal recommendation. Our method follows a four-stage pipeline: (i) Decompose graph-based multimodal signals into spectral bands via graph-guided transformations to isolate semantic granularity; (ii) Modulate band-level reliability with spectral band masking, a training-time masking with representation-consistency objective that suppresses brittle frequency components; (iii) Fuse complementary frequency cues using hyperspectral reasoning with low-rank cross-band interaction; and (iv) Align modality-specific spectral features via contrastive regularization to promote semantic and structural consistency. Experiments on three real-world benchmarks show consistent gains over strong baselines, particularly under sparse and cold-start settings. Additional analyses indicate that structured spectral modeling improves robustness and provides clearer diagnostics of how different bands contribute to performance. The code is available at https://github.com/llm-ml/SSR.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation LearningWei Yang, Rui Zhong, Yiqun Chen, Shixuan Li 等ACM MM 2025 · 被引用 6 次
- TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal RecommendationWei Yang, Rui Zhong, Zihan Lin, Xiaodan Wang 等SIGIR 2026
它引用的顶会 Paper36
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen 等SIGIR 2022 · 被引用 658 次
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He 等SIGIR 2020 · 被引用 621 次
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu 等ACM MM 2021 · 被引用 350 次
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng 等WWW 2023 · 被引用 326 次
相关 Paper
- Frequency-refined Graph Convolution Network with Cross-modal Wavelet Denoising for RecommendationFeiyu Peng, Chaobo He, Junwei Cheng, Huijuan Hu 等ACM MM 2025 · 被引用 6 次
- DIGEST: Dynamic Graph Refinement with Dual Contrastive Semantic Transfer for Multimodal RecommendationXiangyu Sai, Meysam Madadi, Sergio Escalera, Yong XuSIGIR 2026
- Dynamic Spectral Denoising with Global-Context Attention for Multi-Behavior RecommendationMiaomiao Cai, Yunshan Ma, Fangqi Zhu, Junfeng Fang 等KDD 2026
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
- SGP4SR: Separated-Modality Guided User Preference Learning for Multimodal Sequential RecommendationChanghong Li, Zhiqiang Guo, Guohui Li, Zhong Yang 等AAAI 2026
