LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
Rebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang, Tat-Seng Chua
摘要
As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming services. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform their recommendations. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative signals and integrating them into content representations. However, these approaches leave content representation learning out of scope and utilize frozen, pre-trained content models that may not be aligned with the distribution or format of a specific musical setting. Furthermore, even the musical state-of-the-art content modules are either (1) incompatible with the cold-start setting or (2) unable to effectively integrate cross-modal and relational signals. In this paper, we introduce LARP, a multi-modal cold-start playlist continuation model, to effectively overcome these limitations. LARP is a three-stage contrastive learning framework that integrates both multi-modal and relational signals into its learned representations. Our framework uses increasing stages of task-specific abstraction: within-track (language-audio) contrastive loss, track-track contrastive loss, and track-playlist contrastive loss. Experimental results on two publicly available datasets demonstrate the efficacy of LARP over uni-modal and multi-modal models for playlist continuation in a cold-start setting. Finally, this work pioneers the perspective of addressing cold-start recommendation via relational representation learning. Code and dataset are released at: https://github.com/Rsalganik1123/LARP/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fine-tuning Multimodal Large Language Models for Product BundlingXiaohao Liu, Jie Wu, Zhulin Tao, Yunshan Ma 等KDD 2025 · 被引用 3 次
- Discrete Diffusion for Bundle ConstructionTeng Tu, Ai Li, Yunshan Ma, Shuo Xu 等ICLR 2026
它引用的顶会 Paper9
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Contrastive Learning for Cold-Start RecommendationYinwei Wei, Xiang Wang, Qi Li, Liqiang Nie 等ACM MM 2021 · 被引用 321 次
- MAMO: Memory-Augmented Meta-Optimization for Cold-start RecommendationManqing Dong, Feng Yuan, Lina Yao, Xiwei Xu 等KDD 2020 · 被引用 161 次
- Recommendation for New Users and New Items via Randomized Training and Mixture-of-Experts TransformationZiwei Zhu, Shahin Sefati, Parsa Saadatpanah, James CaverleeSIGIR 2020 · 被引用 86 次
相关 Paper
- CMCLRec: Cross-modal Contrastive Learning for User Cold-start Sequential RecommendationXiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang 等SIGIR 2024 · 被引用 56 次
- Contrastive Collaborative Filtering for Cold-Start Item RecommendationZhihui Zhou, Lilin Zhang, Ning YangWWW 2023 · 被引用 91 次
- A Scalable Framework for Automatic Playlist Continuation on Music Streaming ServicesWalid Bendada, Guillaume Salha-Galvan, Thomas Bouabça, Tristan CazenaveSIGIR 2023 · 被引用 15 次
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi 等ICLR 2024 · 被引用 53 次
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang 等ACM MM 2023 · 被引用 62 次
