Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
Chang Xiao, Zixiaofan Yang
摘要
Can you explain how inflation affects interest rates and the broader economy? ... According to the Taylor Rule, monetary policy should adjust the federal funds rate in response to deviations of actual inflation from the target rate and output from potential GDP. (User pauses to think and digest, while the LLM continues generating content) This increase in rates raises the cost of borrowing, reduces consumer spending and business investment, and can lead to slower economic growth ... Waiting User Normal User (Unaware of Change) Satisfied User How do I make olive oil garlic pasta? Sure! I'd be more than happy to help you with that. Olive oil garlic pasta is a wonderfully simple yet flavorful dish that's perfect for just about any occasion. (Streaming paused due to limited computational resource, leaving the user idle and disengaged) ... Can you explain how inflation affects interest rates and the broader economy? How do I make olive oil garlic pasta? ... According to the Taylor Rule, monetary policy should adjust the federal funds rate in response to deviations of actual inflation from the target rate and output from potential GDP. (Streaming is slowed, allowing the user time to digest the content and freeing up computational resources) ... Sure! I'd be more than happy to help you with that. Olive oil garlic pasta is a wonderfully simple yet flavorful dish that's perfect for just about any occasion. To get started, you'll need a few basic ingredients: spaghetti, garlic, extra virgin olive oil, red pepper flakes, salt, and parsley... (Streaming continues smoothly thanks to sufficient computational resources, and proceeds at a faster pace since the content is low in complexity) Normal User Figure 1: (a) In regular LLM streaming, the streaming speed is not associated with the content being delivered. Complex content may be streamed faster than users can comfortably read, resulting in wasted resource. Conversely, simple content may be streamed too slowly, causing users to wait unnecessarily. (b) In our proposed approach, the streaming speed is adapted based on the estimated content complexity. Complex content is slowed down to support user comprehension and optimize resource usage, while simpler content is streamed more quickly to better align with the user's natural reading speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning 等NSDI 2026 · 被引用 29 次
- Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use AgentsQihao Wang, Yue Hu, Mingzhe Lu, Jiayue Wu 等AAAI 2026
- TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive SchedulingJunyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao 等EuroSys 2026
它引用的顶会 Paper9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- GPTCoach: Towards LLM-Based Physical Activity CoachingMatthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio 等CHI 2025 · 被引用 89 次
- Breaking out of the Lab: Mitigating Mind Wandering with Gaze-Based Attention-Aware Technology in ClassroomsStephen Hutt, Kristina Krasich, James R. Brockmole, Sidney K. D'MelloCHI 2021 · 被引用 56 次
- Customizing Emotional Support: How Do Individuals Construct and Interact With LLM-Powered ChatbotsXi Zheng, Zhuoyang Li, Xinning Gui, Yuhan LuoCHI 2025 · 被引用 47 次
- ASHABot: An LLM-Powered Chatbot to Support the Informational Needs of Community Health WorkersPragnya Ramjee, Mehak Chhokar, Bhuvan Sachdeva, Mahendra Meena 等CHI 2025 · 被引用 26 次
相关 Paper
- From Ember to Blaze: Swift Interactive Video Adaptation via Meta-Reinforcement LearningXuedou Xiao, Mingxuan Yan, Yingying Zuo, Boxi Liu 等INFOCOM 2023 · 被引用 14 次
- Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMsQingru Zhang, Chandan Singh, Liyuan Liu, Xiaodong Liu 等ICLR 2024 · 被引用 76 次
- Analyzing the Rapid Generalization of SFT via the Perspective of Attention Head Activation PatternsYang Zhao, Li Du, Xiao Ding, Kai Xiong 等ACL 2025 · 被引用 1 次
- TLadder: QoE-Centric Video Ladder Optimization with Playback Feedback at Billion ScaleZhuqi Li, He Liu, Shenglan Huang, Baojin Geng 等SIGCOMM 2025 · 被引用 2 次
- Towards Event Prediction in Temporal GraphsWenfei Fan, Ruochun Jin, Ping Lu, Chao Tian 等VLDB 2022 · 被引用 19 次
