Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow
Hanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan, Yi Chang, Luxin Yan
摘要
High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically, existing methods mainly introduce event camera to directly fuse the spatiotemporal features between the two modalities. However, this direct fusion is ineffective, since there exists a large gap due to the heterogeneous data representation between frame and event modalities. To address this issue, we explore a common-latent space as an intermediate bridge to mitigate the modality gap. In this work, we propose a novel common spatiotemporal fusion between frame and event modalities for high-dynamic scene optical flow, including visual boundary localization and motion correlation fusion. Specifically, in visual boundary localization, we figure out that frame and event share the similar spatiotemporal gradients, whose similarity distribution is consistent with the extracted boundary distribution. This motivates us to design the common spatiotemporal gradient to constrain the reference boundary localization. In motion correlation fusion, we discover that the frame-based motion possesses spatially dense but temporally discontinuous correlation, while the event-based motion has spatially sparse but temporally continuous correlation. This inspires us to use the reference boundary to guide the complementary motion knowledge fusion between the two modalities. Moreover, common spatiotemporal fusion can not only relieve the cross-modal feature discrepancy, but also make the fusion process interpretable for dense and continuous optical flow. Extensive experiments have been performed to verify the superiority of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and BetterOrest Kupyn, Tetiana Martyniuk, Junru Wu, Zhangyang WangICCV 2019 · 被引用 1,100 次
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li 等ICCV 2021 · 被引用 402 次
- Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale FusionStepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis 等CVPR 2022 · 被引用 126 次
- Learning to Deblur using Light Field Generated and Real Defocus ImagesLingyan Ruan, Bin Chen, Jizhou Li, Miu-Ling LamCVPR 2022 · 被引用 86 次
- ObjectFusion: Multi-modal 3D Object Detection with Object-Centric FusionQi Cai, Yingwei Pan, Ting Yao, Chong-Wah Ngo 等ICCV 2023 · 被引用 71 次
相关 Paper
- Bring Event into RGB and LiDAR: Hierarchical Visual-Motion Fusion for Scene FlowHanyu Zhou, Yi Chang, Zhiwei ShiCVPR 2024 · 被引用 9 次
- x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge SpaceRuishan Guo, Ciyu Ruan, Haoyang Wang, Zihang Gong 等CVPR 2026
- STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic SceneHanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan 等ICCV 2025 · 被引用 1 次
- Exploring the Common Appearance-Boundary Adaptation for Nighttime Optical FlowHanyu Zhou, Yi Chang, Haoyue Liu, Wending Yan 等ICLR 2024 · 被引用 7 次
- RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow EstimationZhexiong Wan, Yuxin Mao, Jing Zhang, Yuchao DaiICCV 2023 · 被引用 35 次
