Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Shunian Chen, Xinyuan Xie, Zheshu Chen, Owen Lee, Liyan Zhao, Zhan Su, Qilin Sun, Benyou Wang
摘要
High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited unimodal or superficial multimodal information. Drawing inspiration from human auditory perception, which adeptly integrates cross-modal cues and performs sophisticated auditory scene analysis, we introduce a novel two-stage automated pipeline. This pipeline first employs specialized pretrained models to extract diverse contextual cues (e.g., speech, music, general sounds, and visual information from associated video). A large language model (LLM) then synthesizes these rich, multimodal inputs to generate detailed and contextaware audio captions. Key contributions of this work include: (1) the proposed scalable method for fine-grained audio caption generation; (2) FusionAudio, a new large-scale dataset comprising 1.2 million such detailed captions, combined with 6 million QA pairs; and (3) enhanced audio models developed using Fu-sionAudio, specifically a CLAP-based audio encoder with superior audio-text alignment and instruction following. This paper paves the way for more nuanced and accurate automated understanding of complex audio environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 被引用 352 次
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning AbilitiesSreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru 等EMNLP 2024 · 被引用 34 次
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 被引用 23 次
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language RepresentationsHang Li, Wenbiao Ding, Yu Kang, Tianqiao Liu 等EMNLP 2021 · 被引用 9 次
相关 Paper
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go 等EMNLP 2025
- VMChill: A Dataset for Fine-Grained Visual-Musical SynergyXiaowei Chi, Zeyue Tian, Jialiang Chen, Wei XueAAAI 2026
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingYifan Yang, Bing Han, Hui Wang, Wei Wang 等ACL 2026 · 被引用 4 次
- FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio PretrainingXiquan Li, Xuenan Xu, Ziyang Ma, Wenxi Chen 等ACL 2026 · 被引用 2 次
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed PerceptionZiyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu 等ICLR 2026 · 被引用 36 次
