AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
Mingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg, Shiyu Zhao, Ying Wu, Manmohan Chandraker
Abstract
Autonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distribution, with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expensive process of continuously curating and annotating data with significant human effort. We propose to leverage recent advances in vision-language and large language models to design an Automatic Data Engine (AIDE) that automatically identifies issues, efficiently curates data, improves the model through auto-labeling, and verifies the model through generation of diverse scenarios. This process operates iteratively, allowing for continuous self-improvement of the model. We further establish a benchmark for open-world detection on AV datasets to comprehensively evaluate various learning paradigms, demonstrating our method's superior performance at a reduced cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66b23032-b2c9-4aac-b62f-6ef0af5d65cbCited by top-tier papers10
- ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work ZonesAnurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong et al.ICCV 2025 · 4 citations
- No Labels, No Problem: Training Visual Reasoners with Multimodal VerifiersDamiano Marsili, Georgia GkioxariICLR 2026 · 3 citations
- AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property EstimationXinbiao Wang, Yuxuan Du, Zihan Lou, Yang Qian et al.NeurIPS 2025 · 1 citation
- Towards Scalable Spatial Intelligence Via 2D-To-3D Data LiftingXingyu Miao, Haoran Duan, Quanhao Qian, Jiuniu Wang et al.ICCV 2025 · 1 citation
- Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge DevicesTianyi Wang, Zichen Wang, Cong Wang, Yuanchao Shu et al.CVPR 2025
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Unbiased Teacher for Semi-Supervised Object DetectionYen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo et al.ICLR 2021 · 603 citations
Related papers
- Lifting Unlabeled Internet-level Data for 3D Scene UnderstandingYixin Chen, Yaowei Zhang, Huangyue Yu, Junchao He et al.CVPR 2026 · 1 citation
- VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy RobustnessQimao Chen, Fang Li, Shaoqing Xu, Zhiyi Lai et al.AAAI 2026 · 2 citations
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous DrivingJianhua Han, Meng Tian, Jiangtong Zhu, Fan He et al.CVPR 2026 · 10 citations
- The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving ModelsRunhao Mao, Hanshi Wang, Yixiang Yang, Qianli Ma et al.CVPR 2026 · 1 citation
- OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous DrivingKota Shimomura, Masaki Nambata, Atsuya Ishikawa, Ryota Mimura et al.ICCV 2025 · 1 citation
