LaneCMKT: Boosting Monocular 3D Lane Detection with Cross-Modal Knowledge Transfer
Runkai Zhao, Heng Wang, Weidong Cai
摘要
Detecting 3D lane lines from monocular images is garnering increasing attention in the Autonomous Driving (AD) area due to its cost-effective edge. However, current monocular image models capture road scenes lacking 3D spatial awareness, which is error-prone to adverse circumstance changes. In this work, we design a novel cross-modal knowledge transfer scheme, namely LaneCMKT, to address this challenge by transferring 3D geometric cues learned from a pre-trained LiDAR model to the image model. Performing on the unified Bird's-Eye-View (BEV) grid, our monocular image model acts as a student network and benefits from the spatial guidance of the 3D LiDAR teacher model over the intermediate feature space. Since LiDAR points and image pixels are intrinsically two different modalities, to facilitate such heterogeneous feature transfer learning at matching levels, we propose a dual-path knowledge transfer mechanism. We divide the feature space into shallow and deep paths where the image student model is prompted to focus on lane-favored geometric cues from the LiDAR teacher model. We conduct extensive experiments and thorough analysis on the large-scale public benchmark OpenLane. Our model achieves notable improvements over the image baseline by 5.3% and the current BEV-driven SoTA method by 2.7% in the F1 score, without introducing any extra computational overhead. We also observe that the 3D abilities grabbed from the teacher model are critical for dealing with complex spatial lane properties from a 2D perspective.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning ModelRui Yu, Xianghang Zhang, Runkai Zhao, Huaicheng Yan 等ICCV 2025 · 被引用 19 次
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu 等ICCV 2025 · 被引用 1 次
相关 Paper
- DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view RepresentationYueru Luo, Shuguang Cui, Zhen LiICLR 2024 · 被引用 15 次
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie 等ICCV 2023 · 被引用 65 次
- GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D UnderstandingJihao Liu, Tai Wang, Boxiao Liu, Qihang Zhang 等ICCV 2023 · 被引用 22 次
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 等ICLR 2023 · 被引用 28 次
- PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature AlignmentZewen Zheng, Xuemin Zhang, Yongqiang Mou, Xiang Gao 等AAAI 2024 · 被引用 26 次
