End-to-End Model Generation with Large Language Models for Adaptive IoT Application Deployment
Zhenyu Wen, Jintao Feng, Nanjie Yao, Di Wu, Cong Wang, Mincheng Wu, Jianbin Qin, Shibo He
Abstract
The deployment of AI-powered applications on resource-constrained edge devices presents a significant software engineering challenge. While Pruning and Neural Architecture Search (NAS) have shown promise in optimizing model efficiency, their application in edge device deployment is often limited by low levels of automation. In this paper, we introduce LLM-based Adaptive Model Generation (LAMDA), a novel framework that tackle the challenge of adaptive AI deployment as an automated model generation problem. LAMDA empowers an LLM to perform end-to-end design, refinement, and optimization of DNNs to meet the specific hardware constraints. Our approach makes two primary contributions. First, we introduce a serialization technique that transforms complex DNN computation graphs into a structured textual representation that makes model architectures comprehensible and manipulable by an LLM. Further, to ground the generation process in real-world hardware constraints, we integrate a feedback-driven optimization loop. This loop leverages an empirical performance model, trained to correlate architectural patterns with on-device latency, enabling the LLM to reason about and optimize for non-functional requirements. To mitigate architectural “hallucinations”, we incorporate context management and validation to ensure valid generation. We evaluate LAMDA through extensive experiments on public benchmarks and real-world edge devices. The results demonstrate that our framework can autonomously generate and adapt DNNs that satisfy deployment-specific accuracy and latency constraints, significantly advancing the state-of-the-art in automated software adaptation for the AI-enabled edge devices.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0e13adb0-2623-41e5-ab0e-de25bdc8ce6bRelated papers
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang et al.MobiCom 2023 · 55 citations
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 32 citations
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao et al.CVPR 2021
- You only search once: on lightweight differentiable architecture search for resource-constrained embedded platformsXiangzhong Luo, Di Liu, Hao Kong, Shuo Huai et al.DAC 2022 · 13 citations
- LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture SearchChengquan Feng, Li Lyna Zhang, Yuanchi Liu, Jiahang Xu et al.NSDI 2024 · 7 citations
