AVHash: Joint Audio-Visual Hashing for Video Retrieval
Yuxiang Zhou, Zhe Sun, Rui Liu, Yong Chen, Dell Zhang
Abstract
Video hashing is a technique of encoding videos into binary vectors, facilitating efficient video storage and high-speed computation. Current approaches to video hashing predominantly utilize sequential frame images to produce semantic binary codes. However, videos encompass not only visual but also audio signals. Therefore, we propose a tri-level Transformer-based audio-visual hashing technique for video retrieval, named AVHash. It first processes audio and visual signals separately using pre-trained AST and ViT large models, and then projects temporal audio and keyframes into a shared latent semantic space using a Transformer encoder. Subsequently, a gated attention mechanism is designed to fuse the paired audio-visual signals in the video, followed by another Transformer encoder leading to the final video representation. The training of this AVHash model is directed by a video-based contrastive loss as well as a semantic alignment regularization term for audio-visual signals. Experimental results show that AVHash significantly outperforms existing video hashing methods in video retrieval tasks. Furthermore, ablation studies reveal that while video hashing based solely on visual signals achieves commendable mAP scores, the incorporation of audio signals can further boost its performance for video retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2a5e6d6c-83bb-4796-b375-d84f64d5c28cCited by top-tier papers3
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video HashingYong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu et al.SIGIR 2025 · 1 citation
- Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash FunctionShuo Yin, Zhiyuan Yin, Yuqing Hou, Rui Liu et al.AAAI 2026
Related papers
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan et al.ACM MM 2024 · 1 citation
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalBoseung Jeong, Jicheol Park, Sungyeon Kim, Suha KwakCVPR 2025
- Unsupervised Similarity-Fusion Transformer Hashing for Multimodal RetrievalZhan Yang, Binghong Chen, Jiajun Tang, Yinan LiACM MM 2025
- Bit-aware Semantic Transformer Hashing for Multi-modal RetrievalWentao Tan, Lei Zhu, Weili Guan, Jingjing Li et al.SIGIR 2022 · 33 citations
