x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space
Ruishan Guo, Ciyu Ruan, Haoyang Wang, Zihang Gong, Jingao Xu, Xinlei Chen
Abstract
Estimating dense 2D optical flow and 3D scene flow is essential for dynamic scene understanding. Recent work combines images, LiDAR, and event data to jointly predict 2D and 3D motion, yet most approaches operate in separate heterogeneous feature spaces. Without a shared latent space that all modalities can align to, these systems rely on multiple modality-specific blocks, leaving cross-sensor mismatches unresolved and making fusion unnecessarily complex. Event cameras naturally provide a spatiotemporal edge signal, which we can treat as an intrinsic edge field to anchor a unified latent representation, termed the Event Edge Space. Building on this idea, we introduce x 2 -Fusion, which reframes multimodal fusion as representation unification: event-derived spatiotemporal edges define an edge-centric homogeneous space, and image and LiDAR features are explicitly aligned in this shared representation. Within this space, we perform reliability-aware adaptive fusion to estimate modality reliability and emphasize stable cues under degradation. We further employ crossdimension contrast learning to tightly couple 2D optical flow with 3D scene flow. Extensive experiments on both synthetic and real benchmarks show that x 2 -Fusion achieves state-of-the-art accuracy under standard conditions and delivers substantial improvements in challenging scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d42f48-2ec7-4002-9dd2-25bfdfa3eb3eBuilds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object DetectionYingwei Li, Adams Wei Yu, Tianjian Meng, Benjamin Caine et al.CVPR 2022 · 508 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
- Edge-Aware Guidance Fusion Network for RGB-Thermal Scene ParsingWujie Zhou, Shaohua Dong, Caie Xu, Yaguan QianAAAI 2022 · 151 citations
Related papers
- RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow EstimationZhexiong Wan, Yuxin Mao, Jing Zhang, Yuchao DaiICCV 2023 · 35 citations
- Bring Event into RGB and LiDAR: Hierarchical Visual-Motion Fusion for Scene FlowHanyu Zhou, Yi Chang, Zhiwei ShiCVPR 2024 · 9 citations
- Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical FlowHanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan et al.CVPR 2025
- Re-coding for Uncertainties: Edge-awareness Semantic Concordance for Resilient Event-RGB SegmentationNan Bao, Yifan Zhao, Lin Zhu, Jia LiNeurIPS 2025 · 1 citation
- ARES: Unifying Asymmetric RGB-Event Stereo for Probabilistic Scene Flow EstimationJie Long Lee, Gim Hee LeeCVPR 2026
