AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video Hashing
Yong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu, Zhouchen Lin, Dell Zhang
Abstract
Existing video hashing techniques for large-scale video retrieval often overlook inherent audio signals, which can potentially compromise retrieval performance. Incorporating both visual and audio signals, however, complicates neural architecture design, rendering the manual crafting of joint audio-visual neural network models challenging. To address this issue, we propose AV-NAS, a method that leverages data-driven Neural Architecture Search (NAS) within a tailored audio-visual network space to automatically discover the optimal video hashing network. Our approach offers: (1) a versatile multi-level semantic architecture based on audio-visual signals, defining a mixed search space encompassing diverse network modules such as MLP, CNN, Transformer, and Mamba, as well as operations like Add, Hadamard, SiLU, LayerNorm, and Skip; (2) a differentiable relaxation of the combinatorial search problem, converting it into a unified differentiable optimization problem which we tackle through our ''coarse search-pruning-finetuning'' strategy. Our experiments on large-scale video datasets show that AV-NAS can discover architectures distinct from expert designs and lead to substantial performance improvements over current state-of-the-art methods including the recently emerged AVHash.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37fa9e5f-d58e-44d3-88a6-29a7e2bad3c7Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 335 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- GLiT: Neural Architecture Search for Global and Local Image TransformerBoyu Chen, Peixia Li, Chuming Li, Baopu Li et al.ICCV 2021 · 100 citations
- TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationYujing Wang, Yaming Yang, Yiren Chen, Jing Bai et al.AAAI 2020 · 66 citations
Related papers
- AVHash: Joint Audio-Visual Hashing for Video RetrievalYuxiang Zhou, Zhe Sun, Rui Liu, Yong Chen et al.ACM MM 2024 · 4 citations
- Searching for Two-Stream Models in Multivariate Space for Video RecognitionXinyu Gong, Heng Wang, Zheng Shou, Matt Feiszli et al.ICCV 2021 · 9 citations
- VONAS: Network Design in Visual Odometry using Neural Architecture SearchXing Cai, Lanqing Zhang, Chengyuan Li, Ge Li et al.ACM MM 2020 · 8 citations
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- Deep Multimodal Neural Architecture SearchZhou Yu, Yuhao Cui, Jun Yu, Meng Wang et al.ACM MM 2020 · 93 citations
