Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
Chang Xiao, Zixiaofan Yang
Abstract
Can you explain how inflation affects interest rates and the broader economy? ... According to the Taylor Rule, monetary policy should adjust the federal funds rate in response to deviations of actual inflation from the target rate and output from potential GDP. (User pauses to think and digest, while the LLM continues generating content) This increase in rates raises the cost of borrowing, reduces consumer spending and business investment, and can lead to slower economic growth ... Waiting User Normal User (Unaware of Change) Satisfied User How do I make olive oil garlic pasta? Sure! I'd be more than happy to help you with that. Olive oil garlic pasta is a wonderfully simple yet flavorful dish that's perfect for just about any occasion. (Streaming paused due to limited computational resource, leaving the user idle and disengaged) ... Can you explain how inflation affects interest rates and the broader economy? How do I make olive oil garlic pasta? ... According to the Taylor Rule, monetary policy should adjust the federal funds rate in response to deviations of actual inflation from the target rate and output from potential GDP. (Streaming is slowed, allowing the user time to digest the content and freeing up computational resources) ... Sure! I'd be more than happy to help you with that. Olive oil garlic pasta is a wonderfully simple yet flavorful dish that's perfect for just about any occasion. To get started, you'll need a few basic ingredients: spaghetti, garlic, extra virgin olive oil, red pepper flakes, salt, and parsley... (Streaming continues smoothly thanks to sufficient computational resources, and proceeds at a faster pace since the content is low in complexity) Normal User Figure 1: (a) In regular LLM streaming, the streaming speed is not associated with the content being delivered. Complex content may be streamed faster than users can comfortably read, resulting in wasted resource. Conversely, simple content may be streamed too slowly, causing users to wait unnecessarily. (b) In our proposed approach, the streaming speed is adapted based on the estimated content complexity. Complex content is slowed down to support user comprehension and optimize resource usage, while simpler content is streamed more quickly to better align with the user's natural reading speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 878ad3b2-76f9-45d2-9bec-3cd692fb4f3dCited by top-tier papers3
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
- Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use AgentsQihao Wang, Yue Hu, Mingzhe Lu, Jiayue Wu et al.AAAI 2026
- TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive SchedulingJunyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao et al.EuroSys 2026
Builds on9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- GPTCoach: Towards LLM-Based Physical Activity CoachingMatthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio et al.CHI 2025 · 89 citations
- Breaking out of the Lab: Mitigating Mind Wandering with Gaze-Based Attention-Aware Technology in ClassroomsStephen Hutt, Kristina Krasich, James R. Brockmole, Sidney K. D'MelloCHI 2021 · 56 citations
- Customizing Emotional Support: How Do Individuals Construct and Interact With LLM-Powered ChatbotsXi Zheng, Zhuoyang Li, Xinning Gui, Yuhan LuoCHI 2025 · 47 citations
- ASHABot: An LLM-Powered Chatbot to Support the Informational Needs of Community Health WorkersPragnya Ramjee, Mehak Chhokar, Bhuvan Sachdeva, Mahendra Meena et al.CHI 2025 · 26 citations
Related papers
- From Ember to Blaze: Swift Interactive Video Adaptation via Meta-Reinforcement LearningXuedou Xiao, Mingxuan Yan, Yingying Zuo, Boxi Liu et al.INFOCOM 2023 · 14 citations
- Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMsQingru Zhang, Chandan Singh, Liyuan Liu, Xiaodong Liu et al.ICLR 2024 · 76 citations
- Analyzing the Rapid Generalization of SFT via the Perspective of Attention Head Activation PatternsYang Zhao, Li Du, Xiao Ding, Kai Xiong et al.ACL 2025 · 1 citation
- TLadder: QoE-Centric Video Ladder Optimization with Playback Feedback at Billion ScaleZhuqi Li, He Liu, Shenglan Huang, Baojin Geng et al.SIGCOMM 2025 · 2 citations
- Towards Event Prediction in Temporal GraphsWenfei Fan, Ruochun Jin, Ping Lu, Chao Tian et al.VLDB 2022 · 19 citations
