Decoupling Scene Perception and Ego Status: A Multi-Context Fusion Approach for Enhanced Generalization in End-to-End Autonomous Driving
Jiacheng Tang, Mingyue Feng, Jiachao Liu, Yaonong Wang, Jian Pu
Abstract
Modular design of planning-oriented autonomous driving has markedly advanced end-to-end systems. However, existing architectures remain constrained by an over-reliance on ego status, hindering generalization and robust scene understanding. We identify the root cause as an inherent design within these architectures that allows ego status to be easily leveraged as a shortcut. Specifically, the premature fusion of ego status in the upstream BEV encoder allows an information flow from this strong prior to dominate the downstream planning module. To address this challenge, we propose AdaptiveAD, an architectural-level solution based on a multi-context fusion strategy. Its core is a dual-branch structure that explicitly decouples scene perception and ego status. One branch performs scene-driven reasoning based on multi-task learning, but with ego status deliberately omitted from the BEV encoder, while the other conducts ego-driven reasoning based solely on the planning task. A scene-aware fusion module then adaptively integrates the complementary decisions from the two branches to form the final planning trajectory. To ensure this decoupling does not compromise multi-task learning, we introduce a path attention mechanism for ego-BEV interaction and add two targeted auxiliary tasks: BEV unidirectional distillation and autoregressive online mapping. Extensive evaluations on the nuScenes dataset demonstrate that AdaptiveAD achieves state-of-the-art open-loop planning performance. Crucially, it significantly mitigates the over-reliance on ego status and exhibits impressive generalization capabilities across diverse scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8359e656-b067-4992-9afb-a2a3ef38eb4aCited by top-tier papers1
Ask how each one uses itBuilds on12
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossXue Yang, Junchi Yan, Qi Ming, Wentao Wang et al.ICML 2021 · 572 citations
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
- DeepInteraction: 3D Object Detection via Modality InteractionZeyu Yang, Jiaqi Chen, Zhenwei Miao, Wei Li et al.NeurIPS 2022 · 268 citations
Related papers
- Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and PlanningBozhou Zhang, Nan Song, Xin Jin, Li ZhangCVPR 2025
- Dualad: Disentangling the Dynamic and Static World for End-to-End DrivingSimon Doll, Niklas Hanselmann, Lukas Schneider, Richard Schulz et al.CVPR 2024
- HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single DecoderYingqi Tang, Zhuoran Xu, Zhaotie Meng, Erkang ChengICCV 2025 · 4 citations
- Planning-oriented Autonomous DrivingYihan Hu, Jiazhi Yang, Li Chen, Keyu Li et al.CVPR 2023
- Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?Zhiqi Li, Zhiding Yu, Shiyi Lan, Jiahan Li et al.CVPR 2024
