LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
Rebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang, Tat-Seng Chua
Abstract
As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming services. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform their recommendations. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative signals and integrating them into content representations. However, these approaches leave content representation learning out of scope and utilize frozen, pre-trained content models that may not be aligned with the distribution or format of a specific musical setting. Furthermore, even the musical state-of-the-art content modules are either (1) incompatible with the cold-start setting or (2) unable to effectively integrate cross-modal and relational signals. In this paper, we introduce LARP, a multi-modal cold-start playlist continuation model, to effectively overcome these limitations. LARP is a three-stage contrastive learning framework that integrates both multi-modal and relational signals into its learned representations. Our framework uses increasing stages of task-specific abstraction: within-track (language-audio) contrastive loss, track-track contrastive loss, and track-playlist contrastive loss. Experimental results on two publicly available datasets demonstrate the efficacy of LARP over uni-modal and multi-modal models for playlist continuation in a cold-start setting. Finally, this work pioneers the perspective of addressing cold-start recommendation via relational representation learning. Code and dataset are released at: https://github.com/Rsalganik1123/LARP/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00bc5fc8-d136-42d5-bf2f-13a678b9d66dCited by top-tier papers2
- Fine-tuning Multimodal Large Language Models for Product BundlingXiaohao Liu, Jie Wu, Zhulin Tao, Yunshan Ma et al.KDD 2025 · 3 citations
- Discrete Diffusion for Bundle ConstructionTeng Tu, Ai Li, Yunshan Ma, Shuo Xu et al.ICLR 2026
Builds on9
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Contrastive Learning for Cold-Start RecommendationYinwei Wei, Xiang Wang, Qi Li, Liqiang Nie et al.ACM MM 2021 · 321 citations
- MAMO: Memory-Augmented Meta-Optimization for Cold-start RecommendationManqing Dong, Feng Yuan, Lina Yao, Xiwei Xu et al.KDD 2020 · 161 citations
- Recommendation for New Users and New Items via Randomized Training and Mixture-of-Experts TransformationZiwei Zhu, Shahin Sefati, Parsa Saadatpanah, James CaverleeSIGIR 2020 · 86 citations
Related papers
- CMCLRec: Cross-modal Contrastive Learning for User Cold-start Sequential RecommendationXiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang et al.SIGIR 2024 · 56 citations
- Contrastive Collaborative Filtering for Cold-Start Item RecommendationZhihui Zhou, Lilin Zhang, Ning YangWWW 2023 · 91 citations
- A Scalable Framework for Automatic Playlist Continuation on Music Streaming ServicesWalid Bendada, Guillaume Salha-Galvan, Thomas Bouabça, Tristan CazenaveSIGIR 2023 · 15 citations
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi et al.ICLR 2024 · 53 citations
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang et al.ACM MM 2023 · 62 citations
