InFi: end-to-end learnable input filter for resource-efficient mobile-centric inference
Mu Yuan, Lan Zhang, Fengxiang He, Xueting Tong, Xiangyang Li
Abstract
Mobile-centric AI applications put forward high requirements for resource-efficiency of model inference. Input filtering is a promising approach to eliminate the redundancy in the input so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1) theoretical filterability of an inference workload to guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2) robust discriminability of feature embedding to allow input filtering to be widely effective for diverse inference tasks and input content. To answer these questions, we first provide a generic formalization of the input filtering problem and theoretically compare the hypothesis complexity of inference models and their input filters to understand the optimization potential of applying input filtering. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. Based on our framework, we design and implement an input filtering system InFi supporting six input modalities. InFi is the first to support text and sensor signal inputs and model partitioning deployments widely adopted by under-resourced mobile systems. Comprehensive evaluations confirm our theoretical results and show that InFi outperforms strong baselines in applicability, accuracy, and efficiency, owing to its generality and end-to-end learnability. InFi can achieve 8.5X throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics app on mobile platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f6fd539-333a-4a65-b591-923d1760074cCited by top-tier papers7
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang et al.MobiCom 2024 · 40 citations
- PacketGame: Multi-Stream Packet Gating for Concurrent Video Inference at ScaleMu Yuan, Lan Zhang, Xuanke You, Xiang-Yang LiSIGCOMM 2023 · 18 citations
- Region-based Content Enhancement for Efficient Video Analytics at the EdgeWeijun Wang, Liang Mi, Shaowei Cen, Haipeng Dai et al.NSDI 2025 · 12 citations
- FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive LearningYifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu et al.USENIX Security 2024 · 6 citations
- Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic SegmentationMingxuan Yan, Yi Wang, Xuedou Xiao, Zhiqing Luo et al.ACM MM 2023 · 3 citations
Builds on8
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang et al.SIGCOMM 2020 · 264 citations
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
Related papers
- Cross-Camera Inference on the Constrained EdgeJingzong Li, Libin Liu, Hong Xu, Shudeng Wu et al.INFOCOM 2023 · 31 citations
- ResMap: Exploiting Sparse Residual Feature Map for Accelerating Cross-Edge Video AnalyticsNing Chen, Shuai Zhang, Sheng Zhang, Yuting Yan et al.INFOCOM 2023 · 11 citations
- Decode-What-Matters: Frame-Level Parallel Generative Decoding to Accelerate Large-Scale Video AnalyticsXiaokun Wang, Yuting Yan, Sheng Zhang, Andong Zhu et al.ACM MM 2025
- Owl: A Pre-and Post-processing Framework for Video Analytics in Low-light SurroundingsRui-Xiao Zhang, Chaoyang Li, Chenglei Wu, Tianchi Huang et al.INFOCOM 2023 · 7 citations
- ENTRO: Tackling the Encoding and Networking Trade-off in Offloaded Video AnalyticsSeyeon Kim, Kyungmin Bin, Donggyu Yang, Sangtae Ha et al.ACM MM 2023 · 6 citations
