HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction
Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, Jingjian Lin
Abstract
Industrial large-scale recommendation models (LRMs) face the challenge of jointly modeling long-range user behavior sequences and heterogeneous non-sequential features under strict efficiency constraints. However, most existing architectures employ a decoupled pipeline: long sequences are first compressed with a query-token based sequence compressor like LONGER, followed by fusion with dense features through token-mixing modules like RankMixer, which thereby limits both the representation capacity and the interaction flexibility. This paper presents HyFormer, a unified hybrid transformer architecture that tightly integrates long-sequence modeling and feature interaction into a single backbone. From the perspective of sequence modeling, we revisit and redesign query tokens in LRMs, and frame the LRM modeling task as an alternating optimization process that integrates two core components: Query Decoding which expands non-sequential features into Global Tokens and performs long sequence decoding over layer-wise key-value representations of long behavioral sequences; and Query Boosting which enhances cross-query and cross-sequence heterogeneous interactions via efficient token mixing. The two complementary mechanisms are performed iteratively to refine semantic representations across layers. Extensive experiments on billion-scale industrial datasets demonstrate that HyFormer consistently outperforms strong LONGER and RankMixer baselines under comparable parameter and FLOPs budgets, while exhibiting superior scaling behavior with increasing parameters and FLOPs. Large-scale online A/B tests in high-traffic production systems further validate its effectiveness, showing significant gains over deployed state-of-the-art models. These results highlight the practicality and scalability of HyFormer as a unified modeling framework for industrial LRMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ca5cfd1-1b30-4ea0-a0c0-60af915cb57cBuilds on5
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie et al.ICML 2024 · 108 citations
- Scaling Sequential Recommendation Models with TransformersPablo Zivic, Hernán Ceferino Vázquez, Jorge SánchezSIGIR 2024 · 25 citations
Related papers
- Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in RecommendationGuoming Li, Shangyu Zhang, Junwei Pan, Wentao Ning et al.KDD 2026 · 2 citations
- Text Is All You Need: Learning Language Representations for Sequential RecommendationJiacheng Li, Ming Wang, Jin Li, Jinmiao Fu et al.KDD 2023 · 134 citations
- HyMiRec: A Hybrid Multi-interest Learning Framework for LLM-based Sequential RecommendationJingyi Zhou, Cheng Chen, Kai Zuo, Manjie Xu et al.WWW 2026 · 2 citations
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li et al.NeurIPS 2025
- Rankformer: A Graph Transformer for Recommendation based on Ranking ObjectiveSirui Chen, Shen Han, Jiawei Chen, Binbin Hu et al.WWW 2025 · 7 citations
