RealtimeTool: Parallel Decoding for Real-Time LLM Function Calling
Xiaoxin Shi, Jiaxin Wan, Linkang Dong, Wei Jiang, Yue Liu, Zengfeng Huang
Abstract
LLM-based function calling enables intelligent agents to interact with external tools and environments, yet autoregressive decoding imposes a fundamental latency bottleneck that limits real-time applications such as embodied intelligence, game AI, and interactive avatars (e.g., 10 Hz control frequency). We observe that function calling differs fundamentally from free-form text generation: structured outputs exhibit substantial token redundancy (delimiters, parameter names) and weak causal dependencies among arguments---two properties that must be exploited jointly to achieve real-time performance. We present RealtimeTool, which introduces special tokens that serve a dual role: compressing low-entropy tokens (4--6× reduction) while acting as mode selectors that enable independent parallel generation of function name and arguments. This synergistic design achieves 3--6× end-to-end speedup (up to 9.6×) with only +8.2% parallelization overhead, while maintaining competitive or improved accuracy across five benchmarks on Qwen-series models (0.5B--14B). With quantization on a consumer-grade GPU, RealtimeTool reaches 61.2 ms P50 latency at 4B scale---enabling 16 Hz real-time control and bridging the gap between LLM function calling and latency-critical real-world deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 285382a3-c81a-40d2-a62f-67631502c5bdBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
Related papers
- UniCompress: Token Compression for Unified Vision-Language Understanding and GenerationZiyao Wang, Chen Chen, Jingtao Li, Weiming Zhuang et al.CVPR 2026 · 1 citation
- AgentVocab: Structure-Aware Vocabulary Adaptation for Efficient LLM AgentsKai Bian, Haosi Mo, Xuebo Liu, Shuangyong Song et al.ICML 2026
- Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMsHao Kang, Qingru Zhang, Han Cai, Weiyuan Xu et al.NeurIPS 2025 · 15 citations
- TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive SchedulingJunyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao et al.EuroSys 2026
- Accelerated Test-Time Scaling with Model-Free Speculative SamplingWoomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh et al.EMNLP 2025
