Video Question Answering Using Language-Guided Deep Compressed-Domain Video Feature
Nayoung Kim, Seong Jong Ha, Je-Won Kang
摘要
Video Question Answering (Video QA) aims to give an answer to the question through semantic reasoning between visual and linguistic information. Recently, handling large amounts of multi-modal video and language information of a video is considered important in the industry. However, the current video QA models use deep features, suffered from significant computational complexity and insufficient representation capability both in training and testing. Existing features are extracted using pre-trained networks after all the frames are decoded, which is not always suitable for video QA tasks. In this paper, we develop a novel deep neural network to provide video QA features obtained from coded video bit-stream to reduce the complexity. The proposed network includes several dedicated deep modules to both the video QA and the video compression system, which is the first attempt at the video QA task. The proposed network is predominantly model-agnostic. It is integrated into the state-of-the-art networks for improved performance without any computationally expensive motion-related deep models. The experimental results demonstrate that the proposed network outperforms the previous studies at lower complexity. https://github.com/Nayoung-Kim-ICP/VQAC
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CaRiNG: Learning Temporal Causal Representation under Non-Invertible Generation ProcessGuangyi Chen, Yifan Shen, Zhenhao Chen, Xiangchen Song 等ICML 2024 · 被引用 22 次
- SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role LabelingJu-Hee Lee, Je-Won KangCVPR 2024
它引用的顶会 Paper1
相关 Paper
- Advancing Video Question Answering with a Multi-modal and Multi-layer Question Enhancement NetworkMeng Liu, Fenglei Zhang, Xin Luo, Fan Liu 等ACM MM 2023 · 被引用 10 次
- Granularity-Adaptive Spatial Evidence Tokenization for Video Question AnsweringHao Jiang, Yang Jin, Zhicheng Sun, Kun Xu 等AAAI 2025 · 被引用 2 次
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan 等ACM MM 2023 · 被引用 5 次
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo 等ICCV 2025 · 被引用 8 次
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li 等AAAI 2020 · 被引用 9 次
