Interpretable Embedding for Ad-Hoc Video Search
Jiaxin Wu, Chong-Wah Ngo
Abstract
Answering query with semantic concepts has long been the mainstream approach for video search. Until recently, its performance is surpassed by concept-free approach, which embeds queries in a joint space as videos. Nevertheless, the embedded features as well as search results are not interpretable, hindering subsequent steps in video browsing and query reformulation. This paper integrates feature embedding and concept interpretation into a neural network for unified dual-task learning. In this way, an embedding is associated with a list of semantic concepts as an interpretation of video content. This paper empirically demonstrates that, by using either the embedding features or concepts, considerable search improvement is attainable on TRECVid benchmarked datasets. Concepts are not only effective in pruning false positive videos, but also highly complementary to concept-free search, leading to large margin of improvement compared to state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
- Multi-Modal Knowledge Hypergraph for Diverse Image RetrievalYawen Zeng, Qin Jin, Tengfei Bao, Wenfeng LiAAAI 2023 · 40 citations
- Learn to Understand Negation in Video RetrievalZiyue Wang, Aozhu Chen, Fan Hu, Xirong LiACM MM 2022 · 12 citations
Related papers
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang et al.SIGIR 2020 · 131 citations
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon et al.CVPR 2024
- Concept Propagation via Attentional Knowledge Graph Reasoning for Video-Text RetrievalSheng Fang, Shuhui Wang, Junbao Zhuo, Qingming Huang et al.ACM MM 2022 · 10 citations
- 3D Concept Grounding on Neural FieldsYining Hong, Yilun Du, Chunru Lin, Josh Tenenbaum et al.NeurIPS 2022 · 24 citations
