MGTA: Multi-scale Graph Tokens Alignment for CTR Prediction via Pre-trained Language Models
Zhongzhen Wu, Yating Ren, Shuochen Li, Huobin Tan
Abstract
Click-through rate (CTR) prediction is a critical task in personalized recommender systems. Existing methods that align collaborative information from conventional CTR models with semantic information from pre-trained language models (PLMs) have demonstrated superior performance compared to approaches relying on a single information source. However, most of them perform alignment at the embedding level, which introduces noise from heterogeneous vector spaces and limits fine-grained semantic mapping. Moreover, these models highly depend on tabular features, thereby limiting their transferability. To address these challenges, we propose to conduct Multi-scale Graph Tokens Alignment (MGTA) for CTR prediction via pre-trained language models, which enables deep cross-modal information alignment while maintaining strong generalizability. Specifically, MGTA first captures multi-scale graph tokens rich in collaborative signals by decoupling and quantizing graph structures based on graph neural networks (GNNs), and then achieves token-level alignment between collaborative signals and semantic knowledge via PLM fine-tuning. To achieve efficient transfer with MGTA, we further introduce the Cross-domain Token Adapter that enables collaborative signals adaptation by mapping graph tokens from the target domain to the source domain, which necessitates only the injection of target-domain semantic knowledge, in turn reducing fine-tuning time. Extensive experiments on three real-world datasets demonstrate the effectiveness of MGTA compared to existing baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Topic Guided Multi-faceted Semantic Disentanglement for CTR predictionFengxin Li, Zhiqian Yin, Hongyan Liu, Jingcai Guo et al.ACM MM 2025
- ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR PredictionJianghao Lin, Bo Chen, Hangyu Wang, Yunjia Xi et al.WWW 2024 · 58 citations
- GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingZiyin Zhang, Hang Yu, Sage Lee, Peng Di et al.ACL 2025 · 11 citations
- LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token EmbeddingsDuo Wang, Yuan Zuo, Fengzhi Li, Junjie WuNeurIPS 2024 · 99 citations
- Learn to Cross-lingual Transfer with Meta Graph Learning Across Heterogeneous LanguagesZheng Li, Mukul Kumar, William Headden, Bing Yin et al.EMNLP 2020 · 26 citations
