AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation
Yong Zhu, Zhenyu Wen, Tao Wang, Zihua Yang, Xiaoli Zhang, Zhen Hong, Bin Qian, Shibo He, Li Ping Qian, Cong Wang
摘要
Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Moby: Empowering 2D Models for Efficient Point Cloud Analytics on the EdgeJingzong Li, Yik Hong Cai, Libin Liu, Yu Mao 等ACM MM 2023 · 被引用 4 次
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 被引用 69 次
- Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on EdgeLinshen Liu, Boyan Su, Junyue Jiang, Guanlin Wu 等ICCV 2025 · 被引用 8 次
- CrossOver: 3D Scene Cross-Modal AlignmentSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath 等CVPR 2025
- SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object DetectionBonan Ding, Jin Xie, Jing Nie, Jiale CaoAAAI 2025 · 被引用 5 次
