ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources
Jason Wu, Yuyang Yuan, Kang Yang, Lance M. Kaplan, Mani Srivastava
Abstract
Multimodal deep learning systems are deployed in dynamic scenarios due to the robustness afforded by multiple sensing modalities. Nevertheless, they struggle with varying compute resource availability (due to multi-tenancy, device heterogeneity, etc.) and fluctuating quality of inputs (from sensor feed corruption, environmental noise, etc.). Statically provisioned multimodal systems cannot adapt when compute resources change over time, while existing dynamic networks struggle with strict compute budgets. Additionally, both systems often neglect the impact of variations in modality quality. Consequently, modalities suffering substantial corruption may needlessly consume resources better allocated towards other modalities. We propose ADMN, a layer-wise Adaptive Depth Multimodal Network capable of tackling both challenges -it adjusts the total number of active layers across all modalities to meet strict compute resource constraints, and continually reallocates layers across input modalities according to their modality quality. Our evaluations showcase ADMN can match the accuracy of state-of-the-art networks while reducing up to 75% of their floating-point operations.
- Mani Srivastava holds concurrent appointments as a Professor of ECE and CS (joint) at the University of California, Los Angeles, and as an Amazon Scholar at Amazon. This paper describes work performed at UCLA and is not associated with Amazon.
Challenges: Although multimodal deep learning systems are generally robust to variable QoI, a key challenge surrounds the computational efficiency. Most state-of-the-art multimodal networks employ static provisioning in which multimodal inputs are processed by a fixed architecture [3, 4] regardless of their individual utility. Consequently, valuable resources may be wasted on low-QoI modalities. Recent work has explored dynamic networks that train policy networks to reduce computation for easy samples [5,6,7,8]. Unfortunately, explicit consideration of highly dynamic QoI variations among modalities has been largely overlooked. We hypothesize that flexibly allocating computational resources among modalities in accordance with each modality's QoI on a per-sample basis can substantially improve model performance in compute-limited settings.
Additionally, current works also neglect the challenge of dynamic compute resource availability. The environments where multimodal systems are most relevant often impose temporally variable but strictly bounded compute budgets. The maximum budget varies with time according to factors such as thermal throttling, energy fluctuations, or multi-tenancy, and cannot be exceeded at any given moment. Neither statically provisioned models nor dynamic networks are equipped to function under such constraints. Most dynamic networks optimize for average-case efficiency, without mechanisms to constrain the worst-case compute beyond the total cost of the full network. The only partially compatible works are those performing model selection with gating networks [6, 8, 9], which can be adapted to varying compute resource availability by training a set of models for each budget. Aside from requiring an unreasonable amount of training resources, it also impedes the standard practice of loading pretrained weights prior to finetuning [10], as there do not exist publicly available pretrained weights for every possible compute budget. We hypothesize that a single network initialized with pretrained weights, which can dynamically adjust its resource usage, offers an effective solution to the challenge of fluctuating compute resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f6462fa-2a89-4fee-a8bf-935f3dcd2fc4Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
Related papers
- Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile SystemsKai Huang, Xiangyu Yin, Heng Huang, Wei GaoMobiCom 2025
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 2 citations
- PATCH: A Plug-in Framework of Non-blocking Inference for Distributed Multimodal SystemJuexing Wang, Guangjing Wang, Xiao Zhang, Li Liu et al.UbiComp 2023 · 9 citations
- Efficient Multimodal Large Language Model via Dynamic KV Cache QuantizationJiahao Fan, Chien-Ming ChenAAAI 2026
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision AssignmentSangwoo Kwon, Seong Hoon Seo, Jae W. Lee, Yeonhong ParkNeurIPS 2025 · 4 citations
