HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat
摘要
Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-specific relevance. To address these challenges, we introduce HierarQ, a task-aware hierarchical Q-Former based framework that sequentially processes frames to bypass the need for frame sampling, while avoiding LLM’s context length limitations. We introduce a lightweight two-stream language-guided feature modulator to incorporate task awareness in video understanding, with the entity stream capturing frame-level object information within a short context and the scene stream identifying their broader interactions over longer period of time. Each stream is supported by dedicated memory banks which enables our proposed Hierarchical Querying transformer (HierarQ) to effectively capture short and long-term context. Extensive evaluations on 10 video benchmarks across video understanding, question answering, and captioning tasks demonstrate HierarQ’s state-of-the-art performance across most datasets, proving its robustness and efficiency for comprehensive video analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- StreamReady: Learning What to Answer and When in Long Streaming VideosShehreen Azad, Vibhav Vineet, Yogesh S. RawatCVPR 2026 · 被引用 19 次
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language ModelPengteng Li, Pinhao Song, Wuyang Li, Huizai Yao 等NeurIPS 2025 · 被引用 13 次
- DisenQ: Disentangling Q-Former for Activity-BiometricsShehreen Azad, Yogesh Singh RawatICCV 2025 · 被引用 4 次
- RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural LanguageSubrata Biswas, Mohammad Nur Hossain Khan, Bashima IslamEMNLP 2025 · 被引用 3 次
- Punching Bag vs. Punching Person: Motion Transferability in VideosRaiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingBo He, Hengduo Li, Young Kyun Jang, Menglin Jia 等CVPR 2024
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 被引用 6 次
- Threading Keyframe with Narratives: MLLMs as Strong Long Video ComprehendersBo Fang, Yuxin Song, Haoyuan Sun, Qiangqiang Wu 等ICLR 2026 · 被引用 13 次
- ∞-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory ConsolidationSaul José Rodrigues dos Santos, António Farinhas, Daniel C. McNamee, André F. T. MartinsICML 2025
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 等AAAI 2026 · 被引用 1 次
