Data-Efficient Playlist Captioning With Musical and Linguistic Knowledge
Giovanni Gabbolini, Romain Hennequin, Elena V. Epure
摘要
Music streaming services feature billions of playlists created by users, professional editors or algorithms. In this content overload scenario, it is crucial to characterise playlists, so that music can be effectively organised and accessed. Playlist titles and descriptions are proposed in natural language either manually by music editors and users or automatically from pre-defined templates. However, the former is time-consuming while the latter is limited by the vocabulary and covered music themes. In this work, we propose PLAYNTELL, a dataefficient multi-modal encoder-decoder model for automatic playlist captioning. Compared to existing music captioning algorithms, PLAYN-TELL leverages also linguistic and musical knowledge to generate correct and thematic captions. We benchmark PLAYNTELL on a new editorial playlists dataset collected from two major music streaming services. PLAYN-TELL yields 2x-3x higher BLEU@4 and CIDEr than state of the art captioning algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li 等CVPR 2022 · 被引用 169 次
- Prefix-Tuning: Optimizing Continuous Prompts for GenerationXiang Lisa Li, Percy LiangACL 2021
- Normalized and Geometry-Aware Self-Attention Network for Image CaptioningLongteng Guo, Jing Liu, Xinxin Zhu, Peng Yao 等CVPR 2020
相关 Paper
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 被引用 3 次
- LLark: A Multimodal Instruction-Following Language Model for MusicJoshua Patrick Gardner, Simon Durand, Daniel Stoller, Rachel M. BittnerICML 2024 · 被引用 34 次
- SleepLM: Natural-Language Intelligence for Human SleepZongzhe Xu, Zitao Shuai, Eideen Mozaffari, Ravi Aysola 等ICML 2026 · 被引用 10 次
- DiffTell: A High-Quality Dataset for Describing Image Manipulation ChangesZonglin Di, Jing Shi, Yifei Fan, Hao Tan 等ICCV 2025 · 被引用 1 次
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 被引用 9 次
