Bridging the Grounding Gap in VideoQA via Typed Memory for Language-based Belief-State Reasoning
Saman Forouzandeh, Wei Peng, Xinghuo Yu, Mahdi Jalili
Abstract
VideoQA models can be accurate yet often fail to align answers with the correct video segments (the grounding gap). We introduce LINGUA (Language-based INference for Grounded Video Understanding Agent), a memory-based agent that performs grounded VideoQA by reasoning in an explicit linguistic belief state. LINGUA uses five mechanisms: (1) event-driven perception (retains 8--12% of frames while preserving 94% of question-relevant events); (2) typed memory for episodic narratives, semantic affordances, and procedural scripts; (3) Belief-Action-Verification loops with postcondition and temporal checks; (4) meta reflection with contrastive refinement; and (5) Bayesian reliability tracking for continual learning without gradient updates. Built with Gemma3-4B (Ollama, 4-bit), LINGUA outperforms strong baselines on five VideoQA benchmarks, reaching 82.4% on NExT-QA and 42.3% Acc@GQA on NExT-GQA (answer + IoU0.5 temporal localization), while running 2.6 faster than dense-frame methods. In continual learning over 100 videos, accuracy rises from 45.2% (first 10) to 61.8% (last 10) without catastrophic forgetting, indicating online adaptation via memory refinement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a5d1686-4f3c-41e5-bcd9-e6aa8bbda228Builds on14
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
Related papers
- MoReVQA: Exploring Modular Reasoning Models for Video Question AnsweringJuhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho et al.CVPR 2024 · 27 citations
- DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question AnsweringBorui Li, Xingcai Zhang, Tianen Liu, Shuai Wang et al.AAAI 2026
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 44 citations
- Empowering Large Language Model for Continual Video Question Answering with Collaborative PromptingChen Cai, Zheng Wang, Jianjun Gao, Wenyang Liu et al.EMNLP 2024 · 3 citations
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video ReasoningYe Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng ShouICLR 2026 · 23 citations
