4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
Chiao-An Yang, Ryo Hachiuma, Sifei Liu, Subhashree Radhakrishnan, Raymond A. Yeh, Yu-Chiang Frank Wang, Min-Hung Chen
摘要
Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal understanding. Existing 3D and 4D Video Question Answering (VQA) benchmarks also emphasize static scenes and lack region-level prompting. We tackle these issues by introducing: (a) 4D-RGPT, a specialized MLLM designed to capture 4D representations from video inputs with enhanced temporal perception; (b) Perceptual 4D Distillation (P4D), a training framework that transfers 4D representations from a frozen expert model into 4D-RGPT for comprehensive 4D perception; and (c) R4D-Bench, a benchmark for depth-aware dynamic scenes with region-level prompting, built via a hybrid automated and human-verified pipeline. Our 4D-RGPT achieves notable improvements on both existing 4D VQA benchmarks and the proposed R4D-Bench benchmark. Links: Project Page | GitHub | Dataset (a) Baseline R4D-Bench (b) Ours Ans: I am not sure.
0 2 12 ⟨R1⟩ Q: What is the average speed of ⟨R1⟩? Frame 1 Frame 2 Frame N () () Ans: 7.0 m/s Region (2D): Where is ⟨R1⟩? Depth (3D): How far away is ⟨R1⟩? Time (4D): When does ⟨R1⟩ move? Track ⟨R1⟩ across frames Depth Perception Temporal Perception Figure 1 | Overview of Region-level 4D Understanding. 4D region-level VQA, e.g., our R4D-Bench, requires MLLMs to be able to track regions (2D), perceive depth (3D), and temporal progression (4D).
Baseline MLLMs cannot recognize one or more of these aspects and thus fail to answer questions correctly. With our distillation framework, our 4D-RGPT better perceives these aspects and answers accurately. We note that the regions labeled with (*) are not provided in R4D-Bench; they are visualized for readability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper43
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- 4D-Bench: Benchmarking Multi-Modal Large Language Models for 4D Object UnderstandingWenxuan Zhu, Bing Li, Cheng Zheng, Jinjie Mai 等ICCV 2025 · 被引用 2 次
- 4DP-QA: Scalable QA for 4D Perception in Vision Language ModelsSeokju Cho, Abhishek Badki, Hang Su, Jindong Jiang 等CVPR 2026 · 被引用 1 次
- VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsShijie Zhou, Alexander Vilesov, Xuehai He, Ziyu Wan 等ICCV 2025 · 被引用 9 次
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu 等CVPR 2026 · 被引用 15 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
