H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
Siran Chen, Yuxiao Luo, Yue Ma, Yu Qiao, Yali Wang
摘要
With the prevalence of Multimodal Large Language Models(MLLMs), autonomous driving has encountered new opportunities and challenges. In particular, multi-modal video understanding is critical to interactively analyze what will happen in the procedure of autonomous driving. However, videos in such a dynamical scene that often contains complex spatial-temporal movements, which restricts the generalization capacity of the existing MLLMs in this field. To bridge the gap, we propose a novel Hierarchical Mamba Adaptation (H-MBA) framework to fit the complicated motion changes in autonomous driving videos. Specifically, our H-MBA consists of two distinct modules, including Context Mamba (C-Mamba) and Query Mamba (Q-Mamba). First, C-Mamba contains various types of structure state space models, which can effectively capture multi-granularity video context for different temporal resolution. Second, Q-Mamba flexibly transforms the current frame as the learnable query, and attentively select multi-granularity video context into query. Consequently, it can adaptively integrate all the video contexts of multi-scale temporal resolutions to enhance video understanding. Via a plug-and-play paradigm in MLLMs, our H-MBA shows the remarkable performance on multi-modal video tasks in autonomous driving, e.g., for risk object detection, it outperforms the previous SOTA method with 5.5% mIoU improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlZeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang 等ICLR 2026 · 被引用 29 次
- HydraMamba: Multi-Head State Space Model for Global Point Cloud LearningKanglin Qu, Pan Gao, Qun Dai, Yuanhao SunACM MM 2025 · 被引用 2 次
- MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingTongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang 等ICCV 2025
- Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud RepresentationKanglin Qu, Pan Gao, Qun Dai, Yuanhao SunICLR 2026
- DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question AnsweringBorui Li, Xingcai Zhang, Tianen Liu, Shuai Wang 等AAAI 2026
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video UnderstandingShehreen Azad, Vibhav Vineet, Yogesh Singh RawatCVPR 2025
- World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingMingliang Zhai, Cheng Li, Zengyuan Guo, Ningrui Yang 等AAAI 2025 · 被引用 9 次
- M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous DrivingXuesong Chen, Shaoshuai Shi, Tao Ma, Jingqiu Zhou 等AAAI 2025 · 被引用 14 次
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and PlanningZhe Liu, Runhui Huang, Rui Yang, Siming Yan 等CVPR 2026 · 被引用 15 次
- When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language ModelsHaicheng Liao, Yongkang Li, Chengyue Wang, Yanchen Guan 等ACM MM 2024 · 被引用 11 次
