InFi: end-to-end learnable input filter for resource-efficient mobile-centric inference
Mu Yuan, Lan Zhang, Fengxiang He, Xueting Tong, Xiangyang Li
摘要
Mobile-centric AI applications put forward high requirements for resource-efficiency of model inference. Input filtering is a promising approach to eliminate the redundancy in the input so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1) theoretical filterability of an inference workload to guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2) robust discriminability of feature embedding to allow input filtering to be widely effective for diverse inference tasks and input content. To answer these questions, we first provide a generic formalization of the input filtering problem and theoretically compare the hypothesis complexity of inference models and their input filters to understand the optimization potential of applying input filtering. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. Based on our framework, we design and implement an input filtering system InFi supporting six input modalities. InFi is the first to support text and sensor signal inputs and model partitioning deployments widely adopted by under-resourced mobile systems. Comprehensive evaluations confirm our theoretical results and show that InFi outperforms strong baselines in applicability, accuracy, and efficiency, owing to its generality and end-to-end learnability. InFi can achieve 8.5X throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics app on mobile platforms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang 等MobiCom 2024 · 被引用 40 次
- PacketGame: Multi-Stream Packet Gating for Concurrent Video Inference at ScaleMu Yuan, Lan Zhang, Xuanke You, Xiang-Yang LiSIGCOMM 2023 · 被引用 18 次
- Region-based Content Enhancement for Efficient Video Analytics at the EdgeWeijun Wang, Liang Mi, Shaowei Cen, Haipeng Dai 等NSDI 2025 · 被引用 12 次
- FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive LearningYifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu 等USENIX Security 2024 · 被引用 6 次
- Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic SegmentationMingxuan Yan, Yi Wang, Xuedou Xiao, Zhiqing Luo 等ACM MM 2023 · 被引用 3 次
它引用的顶会 Paper8
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang 等SIGCOMM 2020 · 被引用 264 次
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu 等ACL 2020 · 被引用 148 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
相关 Paper
- Cross-Camera Inference on the Constrained EdgeJingzong Li, Libin Liu, Hong Xu, Shudeng Wu 等INFOCOM 2023 · 被引用 31 次
- ResMap: Exploiting Sparse Residual Feature Map for Accelerating Cross-Edge Video AnalyticsNing Chen, Shuai Zhang, Sheng Zhang, Yuting Yan 等INFOCOM 2023 · 被引用 11 次
- Decode-What-Matters: Frame-Level Parallel Generative Decoding to Accelerate Large-Scale Video AnalyticsXiaokun Wang, Yuting Yan, Sheng Zhang, Andong Zhu 等ACM MM 2025
- Owl: A Pre-and Post-processing Framework for Video Analytics in Low-light SurroundingsRui-Xiao Zhang, Chaoyang Li, Chenglei Wu, Tianchi Huang 等INFOCOM 2023 · 被引用 7 次
- ENTRO: Tackling the Encoding and Networking Trade-off in Offloaded Video AnalyticsSeyeon Kim, Kyungmin Bin, Donggyu Yang, Sangtae Ha 等ACM MM 2023 · 被引用 6 次
