TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
Zongsheng Cao, Yangfan He, Anran Liu, Jun Xie, Feng Chen, Zhepeng Wang
摘要
Large Video Language Models (LVLMs) have rapidly emerged as the focus of multimedia AI research. Nonetheless, when confronted with lengthy videos, these models struggle: their temporal windows are narrow, and they fail to notice fine-grained semantic shifts that unfold over extended durations. Moreover, mainstream text-based retrieval pipelines, which rely chiefly on surface-level lexical overlap, ignore the rich temporal interdependence among visual, audio, and subtitle channels. To mitigate these limitations, we propose TV-RAG, a training-free architecture that couples temporal alignment with entropy-guided semantics to improve longvideo reasoning. The framework contributes two main mechanisms: (i) a time-decay retrieval module that injects explicit temporal offsets into the similarity computation, thereby ranking text queries according to their true multimedia context; and (ii) an entropyweighted key-frame sampler that selects evenly spaced, informationdense frames, reducing redundancy while preserving representativeness. By weaving these temporal and semantic signals together, TV-RAG realises a dual-level reasoning routine that can be grafted onto any LVLM without re-training or fine-tuning. The resulting system offers a lightweight, budget-friendly upgrade path and consistently surpasses most leading baselines across established longvideo benchmarks such as Video-MME, MLVU, and LongVideoBench, confirming the effectiveness of our model. The code can be found at https://github.com/AI-Researcher-Team/TV-RAG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory MechanismTao Chen, Kun Zhang, Qiong Wu, Xiao Chen 等CVPR 2026 · 被引用 8 次
- RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation SystemsShaobo Li, Yirui Zhou, Yuan Xu, Kevin Chen 等VLDB 2026 · 被引用 3 次
- VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAGHonghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang 等ACL 2026 · 被引用 2 次
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li 等AAAI 2026 · 被引用 1 次
- RaDAR: Relation-aware Diffusion-Asymmetric Graph Contrastive Learning for RecommendationYixuan Huang, Jiawei Chen, Shengfan Zhang, Zongsheng CaoWWW 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
相关 Paper
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin 等NeurIPS 2025 · 被引用 164 次
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingZhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai 等NeurIPS 2025 · 被引用 19 次
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video UnderstandingXiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed ElhoseinyNeurIPS 2025 · 被引用 37 次
- Select Less, Reason More: Prioritizing Evidence Purity for Video ReasoningXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi HuangCVPR 2026 · 被引用 5 次
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video BenchmarkSeng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao 等CVPR 2026
