Mobile Foundation Model as Firmware
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, Shangguang Wang, Mengwei Xu
摘要
In the current AI era, mobile devices such as smartphones are tasked with executing a myriad of deep neural networks (DNNs) locally. It presents a complex landscape, as these models are highly fragmented in terms of architecture, operators, and implementations. Such fragmentation poses significant challenges to the co-optimization of hardware, systems, and algorithms for efficient and scalable mobile AI.
Inspired by the recent groundbreaking progress in large foundation models, this work introduces a novel paradigm for mobile AI, where mobile OS and hardware jointly manage a foundation model that is capable of serving a wide array of mobile AI tasks. This foundation model functions akin to firmware, unmodifiable by apps or the OS, exposed as a system service to Apps. They can invoke this foundation model through a small, offline fine-tuned "adapter" for various downstream tasks. We propose a tangible design of this vision called M4, and prototype it from publicly available pre-trained models. To assess its capability, we also build a comprehensive benchmark consisting of 38 mobile AI tasks and 50 datasets, spanning 5 multimodal inputs. Extensive experiments demonstrate M4's remarkable results: it achieves comparable accuracy in 85% of tasks, offers enhanced scalability regarding storage and memory, and has much simpler operations. In broader terms, this work paves a new way towards efficient and scalable mobile AI in the post-LLM era.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed InferencesMengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li 等USENIX ATC 2024 · 被引用 78 次
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 等USENIX ATC 2025 · 被引用 41 次
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu 等ASPLOS 2025 · 被引用 38 次
- Demystifying Small Language Models for Edge DeploymentZhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi 等ACL 2025 · 被引用 14 次
- Elastic On-Device LLM ServiceWangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang 等MobiCom 2025 · 被引用 5 次
它引用的顶会 Paper44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Analog Foundation ModelsJulian Büchel, Iason Chalas, Giovanni Acampa, An Chen 等NeurIPS 2025 · 被引用 7 次
- MELTing Point: Mobile Evaluation of Language TransformersStefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed HaddadiMobiCom 2024 · 被引用 32 次
- Scaling Law for Large Wireless ModelsZiheng Liu, Jiayi Zhang, Haoyu Wang, Bokai Xu 等AAAI 2026
- CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMsRana Muhammad Shahroz Khan, Zhen Tan, Ruichen Zhang, Hua Wei 等ICLR 2026
- LegoDNN: block-grained scaling of deep neural networks for mobile visionRui Han, Qinglong Zhang, Chi Harold Liu, Guoren Wang 等MobiCom 2021 · 被引用 51 次
