DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation
Fan Zhang, Jinpeng Chen, Tao Wang, Huan Li, Senzhang Wang, Feifei Kou, Ye Ji, Kaimin Wei, Zhenye Yang
Abstract
Out-of-Town (OOT) recommendation aims to provide personalized suggestions for users in unfamiliar cities. However, OOT recommendation faces two fundamental challenges: the difficulty of reasoning across modalities, as preference signals in disparate formats such as images and text are hard to compare; and the preference deviation problem, since a user's resident and tourist preferences often diverge, rendering simple preference transfer ineffective. To address these challenges, we propose Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation (DiMA), a framework for re-ranking Points of Interest (POIs). To tackle the multimodal challenge, DiMA first leverages Multimodal Large Language Models and Large Language Models (LLMs) to transform heterogeneous POI data into unified semantic tags, enabling both cross-modal reasoning and efficient downstream processing. To address preference deviation, a ``teacher'' LLM executes a custom Chain-of-Thought (CoT) process to disentangle resident and tourist preferences from multi-city histories for re-ranking. Finally, a lightweight student model learns this CoT reasoning via Supervised Fine-Tuning and is then refined with Direct Preference Optimization to align with true user choices, with the potential to surpass the teacher. Extensive experiments on a real-world dataset demonstrate that DiMA significantly enhances the performance of baseline models in the OOT recommendation re-ranking task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentChenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang et al.AAAI 2025 · 141 citations
- CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-TuningPeiyuan Liu, Hang Guo, Tao Dai, Naiqi Li et al.AAAI 2025 · 117 citations
- Harnessing Multimodal Large Language Models for Multimodal Sequential RecommendationYuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang et al.AAAI 2025 · 68 citations
Related papers
- Geography-Aware Large Language Models for Next POI RecommendationWei Liu, Zhao Liu, Muzu Xie, Huaijie Zhu et al.ICDE 2026 · 8 citations
- Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher DistillationWeihai Lu, Xiaoxi Cui, Chenke YinSIGIR 2026
- Align³GR: Unified Multi-Level Alignment for LLM-based Generative RecommendationWencai Ye, Mingjie Sun, Shuhang Chen, Wenjin Wu et al.AAAI 2026 · 2 citations
- Bridge the Domains: Large Language Models Enhanced Cross-domain Sequential RecommendationQidong Liu, Xiangyu Zhao, Yejing Wang, Zijian Zhang et al.SIGIR 2025 · 21 citations
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential RecommendationYu Wang, Yonghui Yang, Le Wu, Yi Zhang et al.SIGIR 2026 · 9 citations
