Video Mirror Detection with the Motion-in-Depth Cue
Alex Warren, Ke Xu, Xin Tian, Gary K. L. Tam, Benjamin W. Wah, Rynson W. H. Lau
Abstract
Detecting mirror regions in RGB videos is essential for scene understanding in applications such as scene reconstruction and robotic navigation. Existing video mirror detectors typically rely on cues like inside-outside mirror correspondences and 2D motion inconsistencies. However, these methods often yield noisy or incomplete predictions when confronted with complex real-world video scenes, especially in areas with occlusion or limited visual features and motions. We observe that human perceive and navigate 3D occluded environments with remarkable ease, owing to Motion-in-Depth (MiD) perception. MiD integrates information from visual appearance (image colors and textures), the way objects move around us in 3D space (3D motions), and their relative distance from us (depth) to determine if something is approaching or receding and to support navigation. Motivated by this neuroscience mechanism, we introduce MiD-VMD, the first approach to explicitly model MiD for video mirror detection. MiD-VMD jointly utilizes contrastive 3D motion, depth, and image features through two novel modules based on a combinational QKV transformer architecture. The Motion-in-Depth Attention Learning (MiD-AL) module captures complementary relationships across these modalities with combinatorial attention and enforces a compact encoding to represent global 3D transformations, resulting in more accurate mirror detection and reduced motion artifacts. The Motion-in-Depth Boundary Detection (MiD-BD) module further sharpens mirror boundaries by leveraging cross-modal attention on 3D motion and depth features. Extensive experiments show that MiD-VMD outperforms current SOTAs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- NeRFReN: Neural Radiance Fields with ReflectionsYuan-Chen Guo, Di Kang, Linchao Bao, Yu He et al.CVPR 2022 · 124 citations
- Bi-directional Object-Context Prioritization Learning for Saliency RankingXin Tian, Ke Xu, Xin Yang, Lin Du et al.CVPR 2022 · 33 citations
Related papers
- Effective Video Mirror Detection with Inconsistent Motion CuesAlex Warren, Ke Xu, Jiaying Lin, Gary K. L. Tam et al.CVPR 2024 · 8 citations
- MVGD-Net: A Novel Motion-aware Video Glass Surface Detection MethodYiwei Lu, Hao Huang, Tao YanAAAI 2026
- Learning to Detect Mirrors from Videos via Dual CorrespondencesJiaying Lin, Xin Tan, Rynson W. H. LauCVPR 2023
- Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D MappingAdam W. Harley, Shrinidhi Kowshika Lakshmikanth, Fangyu Li, Xian Zhou et al.ICLR 2020 · 31 citations
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
