HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
Jiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LIN
Abstract
Edge-cloud collaborative inference is crucial for LLM-powered edge devices, as on-device models often lack the required reasoning capability, while cloud-only inference can be costly and slow under strict latency and token/API budgets. However, existing edge-cloud collaboration methods typically route input tasks based on their estimated difficulty. These static, coarse heuristics overlook subtask dependencies, missing opportunities for parallel execution and budget-adaptive routing. To this end, we propose HybridFlow, a resourceadaptive edge-cloud inference framework that enables parallel execution of interdependent subtasks. Specifically, we build a dependency-aware DAG for each input task, facilitating concurrent execution of subtasks once their dependencies are resolved, thereby reducing end-to-end latency. Additionally, we propose a dynamic benefit-cost utility model, optimizing the trade-off between accuracy, token/API cost, and latency in realtime. This dynamic routing minimizes unnecessary cloud usage while preserving reasoning quality. Across GPQA, MMLU-Pro, AIME24, and LiveBench-Reasoning, HybridFlow improves the cost-accuracy trade-off, reducing latency and cloud API usage while maintaining competitive accuracy. Code: https://github.com/ WanyuGroup/ICML2026_HybridFlow
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbed385e-e750-4cf2-bc1f-32e9fe6c2e76Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du et al.ICLR 2024 · 153 citations
Related papers
- AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceMin Li, Yu Li, Ye Tian, Li Jiang et al.DAC 2021 · 37 citations
- Distributed Inference with Deep Learning Models across Heterogeneous Edge DevicesChenghao Hu, Baochun LiINFOCOM 2022 · 81 citations
- JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceZheming Yang, Wen Ji, Qi Guo, Zhi WangACM MM 2023 · 23 citations
- TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank AdaptationZian Wang, Ziyi Wang, Haonan Jin, Jie Xing et al.EuroSys 2026
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 31 citations
