LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
Hanyu Zhou, Gim Hee Lee
摘要
Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suffer from temporal sparsity that limits language-vision temporal coordination. To address this issue, we introduce LLaFEA (Large Language and Frame-Event Assistant) to leverage event cameras for temporally dense perception and frame-event fusion. Our approach employs a cross-attention mechanism to integrate complementary spatial and temporal features, followed by self-attention matching for global spatio-temporal associations. We further embed textual position and duration tokens into the fused visual space to enhance fine-grained alignment. This unified framework ensures robust spatio-temporal coordinate alignment, enabling LMMs to interpret scenes at any position and any time. In addition, we construct a dataset of real-world frames-events with coordinate instructions and conduct extensive experiments to validate the effectiveness of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EventFlash: Towards Efficient MLLMs for Event-Based VisionShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 等ICLR 2026 · 被引用 5 次
- EventDrive: Event Cameras for Vision-Language Driving IntelligenceDongyue Lu, Rong Li, Ao Liang, Lingdong Kong 等CVPR 2026 · 被引用 2 次
- Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language ModelsBaoheng Zhang, Jiahui Liu, Gui Zhao, Weizhou Zhang 等CVPR 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene UnderstandingHanyu Zhou, Gim Hee LeeICLR 2026 · 被引用 20 次
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li 等NeurIPS 2025 · 被引用 7 次
- LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingHongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang 等CVPR 2025
- Flash-Vstream: Efficient Real-Time Understanding for Long Video StreamsHaoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu 等ICCV 2025 · 被引用 15 次
- Spatio-Temporal LLM: Reasoning about Environments and ActionsHaozhen Zheng, beitong tian, Mingyuan Wu, Zhenggang Tang 等ICML 2026 · 被引用 4 次
