TAVT: Towards Transferable Audio-Visual Text Generation
Wang Lin, Tao Jin, Wenwen Pan, Linjun Li, Xize Cheng, Ye Wang, Zhou Zhao
摘要
Audio-visual text generation aims to understand multi-modality contents and translate them into texts. Although various transfer learning techniques of text generation have been proposed, they focused on uni-modal analysis (e.g. text-to-text, visual-to-text) and lack consideration of multi-modal content and cross-modal relation. Motivated by the fact that humans can recognize the timbre of the same low-level concepts (e.g., footstep, rainfall, and laughing), even in different visual conditions, we aim to mitigate the domain discrepancies by audio-visual correlation.In this paper, we propose a novel Transferable Audio-Visual Text Generation framework, named TAVT, which consists of two key components: Audio-Visual Meta-Mapper (AVMM) and Dual Counterfactual Contrastive Learning (DCCL). (1) AVMM first introduces a universal auditory semantic space and drifts the domain-invariant low-level concepts into visual prefixes. Then the reconstruct-based learning encourages the AVMM to learn “which pixels belong to the same sound” and achieve audio-enhanced visual prefix. The well-trained AVMM can be further applied to uni-modal setting. (2) Furthermore, DCCL leverages the destructive counterfactual transformations to provide cross-modal constraints for AVMM from the perspective of feature distribution and text generation. (3) The experimental results show that TAVT outperforms the state-of-the-art methods across multiple domains (cross-datasets, cross-categories) and various modal settings (uni-modal, multi-modal).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Unified Multimodal Editing with Enhanced Knowledge CollaborationKaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu 等NeurIPS 2024 · 被引用 27 次
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li 等ACL 2025 · 被引用 16 次
- Low-rank Prompt Interaction for Continual Vision-Language RetrievalWeicai Yan, Ye Wang, Wang Lin, Zirun Guo 等ACM MM 2024 · 被引用 8 次
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkWang Lin, Feng Wang, Majun Zhang, Wentao Hu 等ICLR 2026 · 被引用 2 次
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather RemovalHanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang 等ACM MM 2025
它引用的顶会 Paper10
- Scaling Up Vision-Language Pretraining for Image CaptioningXiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang 等CVPR 2022 · 被引用 203 次
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Source-Free Domain Adaptation via Distribution EstimationNing Ding, Yixing Xu, Yehui Tang, Chao Xu 等CVPR 2022 · 被引用 134 次
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang 等CVPR 2022 · 被引用 95 次
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 被引用 94 次
相关 Paper
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang 等ICCV 2025 · 被引用 3 次
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 被引用 46 次
- Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation LearningNikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi M. KalayehCVPR 2024
- AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint ControlXinyue Guo, Xiaoran Yang, Lipan Zhang, Jianxuan Yang 等AAAI 2026
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
