Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment
Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, Mohamed Omar
摘要
Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advancement by ECLIPSE has improved long-range text-to-video retrieval by developing an audiovisual video representation. Nonetheless, the objective of the text-to-video retrieval task is to capture the complementary audio and video information that is pertinent to the text query rather than simply achieving better audio and video alignment. To address this issue, we introduce TEFAL, a TExt-conditioned Feature ALignment method that produces both audio and video representations conditioned on the text query. Instead of using only an audiovisual attention block, which could suppress the audio information relevant to the text query, our approach employs two independent cross-modal attention blocks that enable the text to attend to the audio and video representations separately. Our proposed method’s efficacy is demonstrated on four benchmark datasets that include audio: MSR-VTT, LSMDC, VATEX, and Charades, and achieves better than state-of-the-art performance consistently across the four datasets. This is attributed to the additional text-query-conditioned audio representation and the complementary information it adds to the text-query-conditioned video representation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalJiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu 等CVPR 2024 · 被引用 52 次
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan 等NeurIPS 2024 · 被引用 16 次
- Queries Are Not Alone: Clustering Text Embeddings for Video SearchPeiyang Liu, Xi Wang, Ziqiang Cui, Wei YeSIGIR 2025 · 被引用 8 次
- Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video RetrievalYang Liu, Shudong Huang, Deng Xiong, Jiancheng LvAAAI 2025 · 被引用 4 次
- Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalZhiqian Zhao, Liang Li, Lei Shen, Xichun Sheng 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
相关 Paper
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang 等NeurIPS 2025 · 被引用 6 次
- SAVE: Speech-Aware Video Representation Learning for Video-Text RetrievalRuixiang Zhao, Zhihao Xu, Bangxiang Lan, Zijie Xin 等CVPR 2026
- T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video RetrievalYili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu 等ACM MM 2025
- TeachText: CrossModal Generalized Distillation for Text-Video RetrievalIoana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin 等ICCV 2021 · 被引用 147 次
- Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalBoseung Jeong, Jicheol Park, Sungyeon Kim, Suha KwakCVPR 2025
