HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
Jiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LIN
摘要
Edge-cloud collaborative inference is crucial for LLM-powered edge devices, as on-device models often lack the required reasoning capability, while cloud-only inference can be costly and slow under strict latency and token/API budgets. However, existing edge-cloud collaboration methods typically route input tasks based on their estimated difficulty. These static, coarse heuristics overlook subtask dependencies, missing opportunities for parallel execution and budget-adaptive routing. To this end, we propose HybridFlow, a resourceadaptive edge-cloud inference framework that enables parallel execution of interdependent subtasks. Specifically, we build a dependency-aware DAG for each input task, facilitating concurrent execution of subtasks once their dependencies are resolved, thereby reducing end-to-end latency. Additionally, we propose a dynamic benefit-cost utility model, optimizing the trade-off between accuracy, token/API cost, and latency in realtime. This dynamic routing minimizes unnecessary cloud usage while preserving reasoning quality. Across GPQA, MMLU-Pro, AIME24, and LiveBench-Reasoning, HybridFlow improves the cost-accuracy trade-off, reducing latency and cloud API usage while maintaining competitive accuracy. Code: https://github.com/ WanyuGroup/ICML2026_HybridFlow
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du 等ICLR 2024 · 被引用 153 次
相关 Paper
- AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceMin Li, Yu Li, Ye Tian, Li Jiang 等DAC 2021 · 被引用 37 次
- Distributed Inference with Deep Learning Models across Heterogeneous Edge DevicesChenghao Hu, Baochun LiINFOCOM 2022 · 被引用 81 次
- JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceZheming Yang, Wen Ji, Qi Guo, Zhi WangACM MM 2023 · 被引用 23 次
- TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank AdaptationZian Wang, Ziyi Wang, Haonan Jin, Jie Xing 等EuroSys 2026
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 被引用 31 次
