Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, Yu Shi
Abstract
Large-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework ("Generative Recommenders"), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundational models in recommendations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2594cbcf-7c96-4856-92c8-d6f534d939d2Cited by top-tier papers70
- On Softmax Direct Preference Optimization for RecommendationYuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang et al.NeurIPS 2024 · 126 citations
- Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense RepresentationsYuhao Yang, Zhi Ji, Zhaopeng Li, Yi Li et al.NeurIPS 2025 · 90 citations
- Relational Graph TransformerVijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico Lopez et al.ICLR 2026 · 35 citations
- OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce SearchBen Chen, Xian Guo, Siyuan Wang, Zihan Liang et al.ICML 2026 · 23 citations
- Inductive Generative Recommendation via Retrieval-based SpeculationYijie Ding, Jiacheng Li, Julian J. McAuley, Yupeng HouAAAI 2026 · 19 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative RecommendersZhimin Chen, Chenyu Zhao, Ka Chun Mo, Yunjiang Jiang et al.ICLR 2026 · 13 citations
- Scaling Sequential Recommendation Models with TransformersPablo Zivic, Hernán Ceferino Vázquez, Jorge SánchezSIGIR 2024 · 25 citations
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang et al.SC 2020 · 38 citations
- HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR PredictionYunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin et al.SIGIR 2026 · 4 citations
- Scaling Transformers for Discriminative Recommendation via Generative PretrainingChunqi Wang, Bingchao Wu, Zheng Chen, Lei Shen et al.KDD 2025 · 1 citation
