SepFusion: Finding Optimal Fusion Structures for Visual Sound Separation
Dongzhan Zhou, Xinchi Zhou, Di Hu, Hang Zhou, Lei Bai, Ziwei Liu, Wanli Ouyang
摘要
Multiple modalities can provide rich semantic information; and exploiting such information will normally lead to better performance compared with the single-modality counterpart. However, it is not easy to devise an effective cross-modal fusion structure due to the variations of feature dimensions and semantics, especially when the inputs even come from different sensors, as in the field of audio-visual learning. In this work, we propose SepFusion, a novel framework that can smoothly produce optimal fusion structures for visual-sound separation. The framework is composed of two components, namely the model generator and the evaluator. To construct the generator, we devise a lightweight architecture space that can adapt to different input modalities. In this way, we can easily obtain audio-visual fusion structures according to our demands. For the evaluator, we adopt the idea of neural architecture search to select superior networks effectively. This automatic process can significantly save human efforts while achieving competitive performances. Moreover, since our SepFusion provides a series of strong models, we can utilize the model family for broader applications, such as further promoting performance via model assembly, or providing suitable architectures for the separation of certain instrument classes. These potential applications further enhance the competitiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 被引用 39 次
- Boosting Audio Visual Question Answering via Key Semantic-Aware CuesGuangyao Li, Henghui Du, Di HuACM MM 2024 · 被引用 16 次
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 被引用 4 次
- Independency Adversarial Learning for Cross-Modal Sound SeparationZhenkai Lin, Yanli Ji, Yang YangAAAI 2024 · 被引用 1 次
它引用的顶会 Paper16
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 被引用 384 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond ClassificationHang Xu, Lewei Yao, Zhenguo Li, Xiaodan Liang 等ICCV 2019 · 被引用 197 次
相关 Paper
- Multi-Modal and Multi-Scale Temporal Fusion Architecture Search for Audio-Visual Video ParsingJiayi Zhang, Weixin LiACM MM 2023 · 被引用 5 次
- iQuery: Instruments as Queries for Audio-Visual Sound SeparationJiaben Chen, Renrui Zhang, Dongze Lian, Jiaqi Yang 等CVPR 2023
- MFH-NAS:A Hybrid Neural Architecture Search Framework for Multimodal Fusion Object DetectionQuanWei Gao, Shuqi Zhao, Ruyu Wang, Shuyin Zhang 等ICML 2026
- SelM: Selective Mechanism based Audio-Visual SegmentationJiaxu Li, Songsong Yu, Yifan Wang, Lijun Wang 等ACM MM 2024 · 被引用 5 次
- A Unified Audio-Visual Learning Framework for Localization, Separation, and RecognitionShentong Mo, Pedro MorgadoICML 2023 · 被引用 27 次
