TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal Models
Zhenkun Gao, Xuhong Wang, Xin Tan, Yuan Xie
Abstract
Multimodal Large Language Models (MLLMs), particularly smaller, deployable variants, exhibit a critical deficiency in understanding temporal and procedural visual data, a bottleneck hindering their application in real-world embodied AI. This gap is largely caused by a systemic failure in training paradigms, which lack large-scale, procedurally coherent data. To address this problem, we introduce TPRU, a large-scale dataset sourced from diverse embodied scenarios such as robotic manipulation and GUI navigation. TPRU is systematically designed to cultivate temporal reasoning through three complementary tasks: Temporal Reordering, Next-Frame Prediction, and Previous-Frame Review. A key feature is the inclusion of challenging negative samples, compelling models to transition from passive observation to active, cross-modal validation. We leverage TPRU with a reinforcement learning (RL) fine-tuning methodology, specifically targeting the enhancement of resource-efficient models. Experiments show our approach yields dramatic gains: on our manually curated TPRU-Test, the accuracy of TPRU-7B soars from 50.33% to 75.70%, a state-of-the-art result that significantly outperforms vastly larger baselines, including GPT-4o. Crucially, these capabilities generalize effectively, demonstrating substantial improvements on established benchmarks. The codebase is available at https://github.com/Stephen-gzk/TPRU/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b15e205-86d1-44c5-8bc5-20e549a53503Cited by top-tier papers1
Ask how each one uses itBuilds on13
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li et al.ICML 2024 · 184 citations
Related papers
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in RoboticsDongyoung Kim, Sumin Park, Huiwon Jang, Jinwoo Shin et al.NeurIPS 2025 · 29 citations
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic ManipulationJiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang et al.ICLR 2025 · 4 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong et al.ICLR 2026 · 28 citations
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
