Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation
Guoming Li, Shangyu Zhang, Junwei Pan, Wentao Ning, Jin Chen, Gengsheng Xue, Chao Zhou, Shudong Huang, Haijie Gu, Menglin Yang
Abstract
Scaling recommendation models is a central challenge in recommender systems. Recently, RankMixer has emerged as an effective solution, operating on a unified token representation and alternating between token mixing and per-token feedforward networks (P-FFNs) to achieve scalable performance. However, RankMixer suffers from embedding collapse, where learned representations have low effective rank, limiting expressivity and underutilizing the expanded representation space. Through empirical analysis and theoretical insights, we identify rigid token mixing and P-FFN modules as the primary causes of this phenomenon, jointly inducing a damped oscillatory trajectory in effective-rank evolution across layers. To address it, we propose RankElastor, a novel architecture that produces spectrum-robust representations with provable collapse mitigation. RankElastor introduces two components: (i) parameterized full mixing, which enables expressive token mixing with improved spectral robustness; and (ii) GLU-improved P-FFNs, which stabilize representation spectra through GLU-style FFN modules. Extensive experiments on large-scale industrial datasets demonstrate that RankElastor consistently improves recommendation performance, mitigates embedding collapse, and exhibits robust scaling behavior. Code is available at this GitHub repository: https://github.com/vasile-paskardlgm/RankElastor
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47291cb7-7e4c-4c99-80fc-2817110734e0Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
Related papers
- HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR PredictionYunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin et al.SIGIR 2026 · 4 citations
- On the Embedding Collapse when Scaling up Recommendation ModelsXingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen et al.ICML 2024 · 55 citations
- Rankformer: A Graph Transformer for Recommendation based on Ranking ObjectiveSirui Chen, Shen Han, Jiawei Chen, Binbin Hu et al.WWW 2025 · 7 citations
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie et al.ICML 2024 · 108 citations
- Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective ApproachJialei Chen, Yuanbo Xu, Yiheng JiangKDD 2025 · 3 citations
