AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video Hashing
Yong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu, Zhouchen Lin, Dell Zhang
摘要
Existing video hashing techniques for large-scale video retrieval often overlook inherent audio signals, which can potentially compromise retrieval performance. Incorporating both visual and audio signals, however, complicates neural architecture design, rendering the manual crafting of joint audio-visual neural network models challenging. To address this issue, we propose AV-NAS, a method that leverages data-driven Neural Architecture Search (NAS) within a tailored audio-visual network space to automatically discover the optimal video hashing network. Our approach offers: (1) a versatile multi-level semantic architecture based on audio-visual signals, defining a mixed search space encompassing diverse network modules such as MLP, CNN, Transformer, and Mamba, as well as operations like Add, Hadamard, SiLU, LayerNorm, and Skip; (2) a differentiable relaxation of the combinatorial search problem, converting it into a unified differentiable optimization problem which we tackle through our ''coarse search-pruning-finetuning'' strategy. Our experiments on large-scale video datasets show that AV-NAS can discover architectures distinct from expert designs and lead to substantial performance improvements over current state-of-the-art methods including the recently emerged AVHash.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 被引用 335 次
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- GLiT: Neural Architecture Search for Global and Local Image TransformerBoyu Chen, Peixia Li, Chuming Li, Baopu Li 等ICCV 2021 · 被引用 100 次
- TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationYujing Wang, Yaming Yang, Yiren Chen, Jing Bai 等AAAI 2020 · 被引用 66 次
相关 Paper
- AVHash: Joint Audio-Visual Hashing for Video RetrievalYuxiang Zhou, Zhe Sun, Rui Liu, Yong Chen 等ACM MM 2024 · 被引用 4 次
- Searching for Two-Stream Models in Multivariate Space for Video RecognitionXinyu Gong, Heng Wang, Zheng Shou, Matt Feiszli 等ICCV 2021 · 被引用 9 次
- VONAS: Network Design in Visual Odometry using Neural Architecture SearchXing Cai, Lanqing Zhang, Chengyuan Li, Ge Li 等ACM MM 2020 · 被引用 8 次
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang 等AAAI 2025 · 被引用 7 次
- Deep Multimodal Neural Architecture SearchZhou Yu, Yuhao Cui, Jun Yu, Meng Wang 等ACM MM 2020 · 被引用 93 次
