PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching
Yun Wang, Junjie Hu, Qiaole Dong, Yongjian Zhang, Yanwei Fu, Tin Lun Lam, Dapeng Wu
摘要
Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users. Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal consistency in a computationally efficient manner. Previous methods attempt to address this by aggregating spatio-temporal information but face a fundamental trade-off: limited temporal modeling provides only modest gains, whereas capturing long-range dependencies significantly increases computational cost. To address this limitation, we introduce a memory buffer for modeling long-range spatio-temporal consistency while achieving efficient dynamic stereo matching. Inspired by the two-stage decision-making process in humans, we propose a Pick-and-Play Memory (PPM) construction module for dynamic Stereo matching, dubbed as PPMStereo. PPM consists of a pick'process that identifies the most relevant frames and a play'process that weights the selected frames adaptively for spatio-temporal aggregation. This two-stage collaborative process maintains a compact yet highly informative memory buffer while achieving temporally consistent information aggregation. Extensive experiments validate the effectiveness of PPMStereo, demonstrating state-of-the-art performance in both accuracy and temporal consistency. % Notably, PPMStereo achieves 0.62/1.11 TEPE on the Sintel clean/final (17.3% &9.02% improvements over BiDAStereo) with fewer computational costs. Codes are available at bluehttps://github.com/cocowy1/PPMStereo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai 等CVPR 2022 · 被引用 294 次
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 等CVPR 2024 · 被引用 95 次
相关 Paper
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova 等CVPR 2023
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu 等ICCV 2019 · 被引用 137 次
- Occupancy Learning with Spatiotemporal MemoryZiyang Leng, Jiawei Yang, Wenlong Yi, Bolei ZhouICCV 2025 · 被引用 10 次
- Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsRuikang Xu, Zeyu Xiao, Mingde Yao, Yueyi Zhang 等ACM MM 2021 · 被引用 20 次
- DeepVideoMVS: Multi-View Stereo on Video With Recurrent Spatio-Temporal FusionArda Düzçeker, Silvano Galliani, Christoph Vogel, Pablo Speciale 等CVPR 2021
