Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
Ruiyu Mao, Baoming Zhang, Nicholas Ruozzi, Yunhui Guo
Abstract
Roadside perception datasets are typically constructed via cooperative labeling between synchronized vehicle and roadside frame pairs. However, real deployment often requires annotation of roadside-only data due to hardware and privacy constraints. Even human experts struggle to produce accurate labels without vehicle-side data (image, LIDAR), which not only increases annotation difficulty and cost, but also reveals a fundamental learnability problem: many roadside-only scenes contain distant, blurred, or occluded objects whose 3D properties are ambiguous from a single view and can only be reliably annotated by cross-checking paired vehicle--roadside frames. We refer to such cases as inherently ambiguous samples. To reduce wasted annotation effort on inherently ambiguous samples while still obtaining high-performing models, we turn to active learning. This work focuses on active learning for roadside monocular 3D object detection and proposes a learnability-driven framework that selects scenes which are both informative and reliably labelable, suppressing inherently ambiguous samples while ensuring coverage. Experiments demonstrate that our method, LH3D, achieves 86.06%, 67.32%, and 78.67% of full-performance for vehicles, pedestrians, and cyclists respectively, using only 25% of the annotation budget on DAIR-V2X-I, significantly outperforming uncertainty-based baselines. This confirms that learnability, not uncertainty, matters for roadside 3D perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d73c1772-a889-4876-b7d7-827edb86a4a5Builds on12
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo et al.CVPR 2022 · 475 citations
- Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection TaskXiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi et al.CVPR 2022 · 130 citations
Related papers
- Active Learning for Lane Detection: A Knowledge Distillation ApproachFengchao Peng, Chao Wang, Jianzhuang Liu, Zhen YangICCV 2021 · 18 citations
- Exploring Active 3D Object Detection from a Generalization PerspectiveYadan Luo, Zhuoxiao Chen, Zijian Wang, Xin Yu et al.ICLR 2023 · 3 citations
- RCooper: A Real-world Large-scale Dataset for Roadside Cooperative PerceptionRuiyang Hao, Siqi Fan, Yingru Dai, Zhenlin Zhang et al.CVPR 2024
- Improving Distant 3D Object Detection Using 2D Box SupervisionZetong Yang, Zhiding Yu, Christopher B. Choy, Renhao Wang et al.CVPR 2024 · 7 citations
- RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground SimulationYuwen Du, Anning Hu, Zichen Chao, Yifan Lu et al.ICCV 2025 · 2 citations
