Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
Luoyi Sun, Xuenan Xu, Mengyue Wu, Weidi Xie
Abstract
Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the following aspects: insufficient volume, simplistic content, and arduous collection procedures. To establish an audio dataset with high-quality captions, we propose an innovative, automatic approach leveraging multimodal inputs, such as video frames, audio streams. Specifically, we construct a large-scale, high-quality, audio-language dataset, named as Auto-ACD, comprising over 1.5M audio-text pairs. We exploit a series of pre-trained models or APIs, to determine audio-visual synchronisation, generate image captions, object detection, or audio tags for specific videos. Subsequently, we employ LLM to paraphrase a congruent caption for each audio, guided by the extracted multi-modality clues. To demonstrate the effectiveness of the proposed dataset, we train widely used models on our dataset and show performance improvement on various downstream tasks, for example, audio-language retrieval, audio captioning, zero-shot classification. In addition, we establish a novel benchmark with environmental information and provide a benchmark for audio-text tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db9e013d-84dc-49de-ba44-6303036602d0Cited by top-tier papers13
- AudSemThinker: Enhancing Audio-Language Models Through Reasoning over Semantics of SoundGijs Wijngaard, Elia Formisano, Michele Esposito, Michel DumontierNeurIPS 2025 · 22 citations
- Omni2Sound: Towards Unified Video-Text-to-Audio Generationyusheng dai, Zehua Chen, Yuxuan Jiang, Qiuhong Ke et al.CVPR 2026 · 12 citations
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding TasksYadong Niu, TIANZI WANG, Heinrich Dinkel, Xingwei Sun et al.ICML 2026 · 11 citations
- AudioStory: Generating Long-Form Narrative Audio with Large Language ModelsYuxin Guo, Teng Wang, Yuying Ge, Shijie Ma et al.CVPR 2026 · 5 citations
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationKaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang JiangICCV 2025 · 3 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
Related papers
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matchingHugo Malard, Michel Olvera, Stéphane Lathuilière, Slim EssidNeurIPS 2024 · 3 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- Towards Fine-grained Audio Captioning with Multimodal Contextual FusionShunian Chen, Xinyuan Xie, Zheshu Chen, Owen Lee et al.ACL 2026
- SEAR: Semantically-grounded Audio RepresentationsRajat Hebbar, Digbalay Bose, Shrikanth NarayananACM MM 2023
