Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Gleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev, Erik Schultheis, Vage Egiazarian, Anton Sinitsin, Denis Kuznedelev, Dan Alistarh
Abstract
Large Language Models (LLMs) have demonstrated the ability to tackle increasingly complex tasks through advanced reasoning, long-form content generation, and tool use. Solving these tasks often involves long inference-time computations. In human problem solving, a common strategy to expedite work is collaboration: by dividing the problem into sub-tasks, exploring different strategies concurrently, etc. Recent research has shown that LLMs can also operate in parallel by implementing explicit cooperation frameworks, such as voting mechanisms or the explicit creation of independent sub-tasks that can be executed in parallel. However, each of these frameworks may not be suitable for all types of tasks, which can hinder their applicability. In this work, we propose a different design approach: we run LLM "workers" in parallel , allowing them to synchronize via a concurrently-updated attention cache and prompt these workers to decide how best to collaborate. Our approach allows the LLM instances to come up with their own collaboration strategy for the problem at hand, all the while "seeing" each other's memory in the concurrent KV cache. We implement this approach via Hogwild! Inference: a parallel LLM inference engine where multiple instances of the same LLM run in parallel with the same attention cache, with "instant" access to each other's memory. 1 Hogwild! Inference takes advantage of Rotary Position Embeddings (RoPE) to avoid recomputation while improving parallel hardware utilization. We find that modern reasoning-capable LLMs can perform inference with shared Key-Value cache out of the box, without additional fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa6ba91c-dde9-4aba-a174-794826f7fcffCited by top-tier papers8
- Parallel-R1: Towards Parallel Thinking via Reinforcement LearningTong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang et al.ICLR 2026 · 53 citations
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot ManipulationLetian Fu, Justin Yu, Karim El-Refai, Ethan Kou et al.ICML 2026 · 36 citations
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen et al.NeurIPS 2025 · 36 citations
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning ModelsEmil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza et al.NeurIPS 2025 · 9 citations
- Planned DiffusionDaniel Mingyi Israel, Tian Jin, Ellie Y Cheng, Guy Van den Broeck et al.ICLR 2026 · 8 citations
Builds on31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Efficient Cooperation-Aware Key and Value Management for LLM InferenceQiheng Sun, Hongwei Zhang, Junxu Liu, Haocheng Xia et al.VLDB 2026
- KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache GenerationMinsik Cho, Mohammad Rastegari, Devang NaikICML 2024 · 12 citations
- Deliberation in Latent Space via Differentiable Cache AugmentationLuyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie et al.ICML 2025
- Synergistic Weak-Strong Collaboration by Aligning PreferencesYizhu Jiao, Xuchao Zhang, Zhaoyang Wang, Yubo Ma et al.ACL 2025
- SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesRuslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen et al.NeurIPS 2024 · 70 citations
