ALCAP: Alignment-Augmented Music Captioner
Zihao He, Weituo Hao, Wei Tsung Lu, Changyou Chen, Kristina Lerman, Xuchen Song
摘要
Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms. Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently ignoring the intricate interplay between the two. However, a comprehensive understanding of music necessitates the integration of both these elements. In this study, we delve into this overlooked realm by introducing a method to systematically learn multimodal alignment between audio and lyrics through contrastive learning. This not only recognizes and emphasizes the synergy between audio and lyrics but also paves the way for models to achieve deeper cross-modal coherence, thereby producing high-quality captions. We provide both theoretical and empirical results demonstrating the advantage of the proposed method, which achieves new state-of-the-art on two music captioning datasets. Our code is publicly available at https://github.com/ zihaohe123/ALCAP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Retrieval-Augmented Reinforcement LearningAnirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber 等ICML 2022 · 被引用 69 次
- A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingHe Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma 等ICML 2022 · 被引用 64 次
相关 Paper
- Accommodating Audio Modality in CLIP for Multimodal ProcessingLudan Ruan, Anwen Hu, Yuqing Song, Liang Zhang 等AAAI 2023 · 被引用 18 次
- An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matchingHugo Malard, Michel Olvera, Stéphane Lathuilière, Slim EssidNeurIPS 2024 · 被引用 3 次
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingYiming Li, Zhifang Guo, Xiangdong Wang, Hong LiuACM MM 2024 · 被引用 9 次
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 被引用 65 次
- Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionThong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai 等EMNLP 2022 · 被引用 8 次
