CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
Juntae Lee, Jihwan Bang, Seunghan Yang, Simyung Chang
摘要
We present CIFLEX (Contextual Instruction Flow for Sub-task Execution), which is a novel execution system for efficient sub-task handling in multi-turn interactions with a single on-device large language model (LLM). As LLMs become increasingly capable, a single model is expected to handle diverse sub-tasks that more effectively and comprehensively support answering user requests. Naive approach reprocesses the entire conversation context when switching between main and sub-tasks (e.g., query rewriting, summarization), incurring significant computational overhead. CIFLEX mitigates this overhead by reusing the key-value (KV) cache from the main task and injecting only task-specific instructions into isolated side paths. After sub-task execution, the model rolls back to the main path via cached context, thereby avoiding redundant prefill computation. To support sub-task selection, we also develop a hierarchical classification strategy tailored for small-scale models, decomposing multi-choice decisions into binary ones. Experiments show that CIFLEX significantly reduces computational costs without degrading task performance, enabling scalable and efficient multi-task dialogue on-device.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao 等NeurIPS 2025 · 被引用 83 次
- Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational AnswersTianhua Zhang, Kun Li, Hongyin Luo, Xixin Wu 等EMNLP 2024 · 被引用 4 次
相关 Paper
- Accelerating LLM Serving for Multi-turn Dialogues with Efficient Resource ManagementJinwoo Jeong, Jeongseob AhnASPLOS 2025 · 被引用 6 次
- Elastic On-Device LLM ServiceWangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang 等MobiCom 2025 · 被引用 5 次
- Chain-of-Instructions: Compositional Instruction Tuning on Large Language ModelsShirley Anugrah Hayati, Taehee Jung, Tristan Bodding-Long, Sudipta Kar 等AAAI 2025 · 被引用 14 次
- Stateful Large Language Model Serving with PensieveLingfan Yu, Jinkun Lin, Jinyang LiEuroSys 2025 · 被引用 23 次
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang 等USENIX ATC 2024 · 被引用 273 次
