PATCH: A Plug-in Framework of Non-blocking Inference for Distributed Multimodal System
Juexing Wang, Guangjing Wang, Xiao Zhang, Li Liu, Huacheng Zeng, Li Xiao, Zhichao Cao, Lin Gu, Tianxing Li
摘要
Recent advancements in deep learning have shown that multimodal inference can be particularly useful in tasks like autonomous driving, human health, and production line monitoring. However, deploying state-of-the-art multimodal models in distributed IoT systems poses unique challenges since the sensor data from low-cost edge devices can get corrupted, lost, or delayed before reaching the cloud. These problems are magnified in the presence of asymmetric data generation rates from different sensor modalities, wireless network dynamics, or unpredictable sensor behavior, leading to either increased latency or degradation in inference accuracy, which could affect the normal operation of the system with severe consequences like human injury or car accident. In this paper, we propose PATCH, a framework of speculative inference to adapt to these complex scenarios. PATCH serves as a plug-in module in the existing multimodal models, and it enables speculative inference of these off-the-shelf deep learning models. PATCH consists of 1) a Masked-AutoEncoder-based cross-modality imputation module to impute missing data using partially-available sensor data, 2) a lightweight feature pair ranking module that effectively limits the searching space for the optimal imputation configuration with low computation overhead, and 3) a data alignment module that aligns multimodal heterogeneous data streams without using accurate timestamp or external synchronization mechanisms. We implement PATCH in nine popular multimodal models using five public datasets and one self-collected dataset. The experimental results show that PATCH achieves up to 13% mean accuracy improvement over the state-of-art method while only using 10% of training data and reducing the training overhead by 73% compared to the original cost of retraining the model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CLAP: Collaborative Adaptation for Patchwork LearningSen Cui, Abudukelimu Wuerkaixi, Weishen Pan, Jian Liang 等ICLR 2024 · 被引用 3 次
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 被引用 13 次
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the EdgeYu Zhang, Xi Zhang, Hualin zhou, Xinyuan Chen 等ICML 2026
- Towards Unified Vision-Language Models with Incomplete Multi-Modal InputsXiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang 等AAAI 2026 · 被引用 1 次
- Missing Value Imputation for Multi-attribute Sensor Data Streams via Message PropagationXiao Li, Huan Li, Hua Lu, Christian S. Jensen 等VLDB 2024 · 被引用 17 次
