Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Shunian Chen, Xinyuan Xie, Zheshu Chen, Owen Lee, Liyan Zhao, Zhan Su, Qilin Sun, Benyou Wang
Abstract
High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited unimodal or superficial multimodal information. Drawing inspiration from human auditory perception, which adeptly integrates cross-modal cues and performs sophisticated auditory scene analysis, we introduce a novel two-stage automated pipeline. This pipeline first employs specialized pretrained models to extract diverse contextual cues (e.g., speech, music, general sounds, and visual information from associated video). A large language model (LLM) then synthesizes these rich, multimodal inputs to generate detailed and contextaware audio captions. Key contributions of this work include: (1) the proposed scalable method for fine-grained audio caption generation; (2) FusionAudio, a new large-scale dataset comprising 1.2 million such detailed captions, combined with 6 million QA pairs; and (3) enhanced audio models developed using Fu-sionAudio, specifically a CLAP-based audio encoder with superior audio-text alignment and instruction following. This paper paves the way for more nuanced and accurate automated understanding of complex audio environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43320daf-25aa-4a2d-a621-b1bb7fb85921Builds on6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 352 citations
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning AbilitiesSreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru et al.EMNLP 2024 · 34 citations
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language RepresentationsHang Li, Wenbiao Ding, Yu Kang, Tianqiao Liu et al.EMNLP 2021 · 9 citations
Related papers
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- VMChill: A Dataset for Fine-Grained Visual-Musical SynergyXiaowei Chi, Zeyue Tian, Jialiang Chen, Wei XueAAAI 2026
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingYifan Yang, Bing Han, Hui Wang, Wei Wang et al.ACL 2026 · 4 citations
- FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio PretrainingXiquan Li, Xuenan Xu, Ziyang Ma, Wenxi Chen et al.ACL 2026 · 2 citations
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed PerceptionZiyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu et al.ICLR 2026 · 36 citations
