Smartadapt: Multi-branch Object Detection Framework for Videos on Mobiles
Ran Xu, Fangzhou Mu, Jayoung Lee, Preeti Mukherjee, Somali Chaterji, Saurabh Bagchi, Yin Li
摘要
Several recent works seek to create lightweight deep net-works for video object detection on mobiles. We observe that many existing detectors, previously deemed computationally costly for mobiles, intrinsically support adaptive inference, and offer a multi-branch object detection frame-work (MBODF). Here, an MBODF is referred to as a so-lution that has many execution branches and one can dy-namically choose from among them at inference time to sat-isfy varying latency requirements (e.g. by varying resolution of an input frame). In this paper, we ask, and answer, the wide-ranging question across all MBODFs: How to expose the right set of execution branches and then how to sched-ule the optimal one at inference time? In addition, we un-cover the importance of making a content-aware decision on which branch to run, as the optimal one is conditioned on the video content. Finally, we explore a content-aware scheduler, an Oracle one, and then a practical one, leveraging various lightweight feature extractors. Our evaluation shows that layered on Faster R-CNN-based MBODF, compared to 7 baselines, our Smartadapt achieves a higher Pareto optimal curve in the accuracy-vs-latency space for the ILSVRC VID dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Eventful Transformers: Leveraging Temporal Redundancy in Vision TransformersMatthew Dutson, Yin Li, Mohit GuptaICCV 2023 · 被引用 18 次
- Learning to Inference Adaptively for Multimodal Large Language ModelsZhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi 等ICCV 2025 · 被引用 4 次
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 被引用 1 次
- Sketch Down the FLOPs: Towards Efficient Networks for Human SketchAneeshan Sain, Subhajit Maity, Pinaki Nath Chowdhury, Subhadeep Koley 等CVPR 2025
它引用的顶会 Paper16
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 236 次
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNetsKai Han, Yunhe Wang, Qiulin Zhang, Wei Zhang 等NeurIPS 2020 · 被引用 115 次
- MicroNet: Improving Image Recognition with Extremely Low FLOPsYunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 等ICCV 2021 · 被引用 108 次
相关 Paper
- LiteReconfig: cost and content aware reconfiguration of video object detection systems for mobile GPUsRan Xu, Jayoung Lee, Pengcheng Wang, Saurabh Bagchi 等EuroSys 2022 · 被引用 24 次
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu 等MobiCom 2021 · 被引用 103 次
- VisFlow: Adaptive Content-Aware Video Analytics on Collaborative CamerasYuting Yan, Sheng Zhang, Xiaokun Wang, Ning Chen 等INFOCOM 2024 · 被引用 10 次
- MobileDets: Searching for Object Detection Architectures for Mobile AcceleratorsYunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin 等CVPR 2021
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
