Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
Zhimin Chen, Chenyu Zhao, Ka Chun Mo, Yunjiang Jiang, Jane H. Lee, Khushhall Chandra Mahajan, Ning Jiang, Kai Ren, Jinhui Li, Wen-Yun Yang
摘要
Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer-like architectures, has led to significant advancements recently (e.g., HSTU, SIM, and TWIN models). While scaling to ultra-long user histories (10k to 100k items) generally improves model performance, it also creates significant challenges on latency, queries per second (QPS) and GPU cost in industry-scale recommendation systems. Existing models do not adequately address these industrial scalability issues. In this paper, we propose a novel two-stage modeling framework, namely VIrtual Sequential Target Attention (VISTA), which decomposes traditional target attention from a candidate item to user history items into two distinct stages: (1) user history summarization into a few hundred tokens; followed by (2) candidate item attention to those tokens. These summarization token embeddings are then cached in storage system and then utilized as sequence features for downstream model training and inference. This novel design for scalability enables VISTA to scale to lifelong user histories (up to one million items) while keeping downstream training and inference costs fixed, which is essential in industry. Our approach achieves significant improvements in offline and online metrics and has been successfully deployed on an industry leading recommendation platform serving billions of users.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song 等ICCV 2023 · 被引用 358 次
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 等ICML 2024 · 被引用 200 次
相关 Paper
- Scaling Sequential Recommendation Models with TransformersPablo Zivic, Hernán Ceferino Vázquez, Jorge SánchezSIGIR 2024 · 被引用 25 次
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 等AAAI 2026 · 被引用 1 次
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic TokenizationGuanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin 等AAAI 2025 · 被引用 18 次
- A Training-Free Sub-quadratic Cost Transformer Model Serving Framework with Hierarchically Pruned AttentionHeejun Lee, Geon Park, Youngwan Lee, Jaduk Suh 等ICLR 2025
- HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR PredictionYunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin 等SIGIR 2026 · 被引用 4 次
