AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation
Yong Zhu, Zhenyu Wen, Tao Wang, Zihua Yang, Xiaoli Zhang, Zhen Hong, Bin Qian, Shibo He, Li Ping Qian, Cong Wang
Abstract
Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e3a7fb57-6111-4922-b0b0-c70014020046Related papers
- Moby: Empowering 2D Models for Efficient Point Cloud Analytics on the EdgeJingzong Li, Yik Hong Cai, Libin Liu, Yu Mao et al.ACM MM 2023 · 4 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
- Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on EdgeLinshen Liu, Boyan Su, Junyue Jiang, Guanlin Wu et al.ICCV 2025 · 8 citations
- CrossOver: 3D Scene Cross-Modal AlignmentSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath et al.CVPR 2025
- SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object DetectionBonan Ding, Jin Xie, Jing Nie, Jiale CaoAAAI 2025 · 5 citations
