Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong, Seungone Kim
Abstract
Large language models (LLMs) are typically prompted to follow a single instruction per inference call. In this work, we analyze whether LLMs also hold the capability to handle multiple instructions simultaneously, denoted as MULTI-TASK INFERENCE. For this purpose, we introduce the MTI BENCH (Multi-Task Inference Benchmark), a comprehensive evaluation benchmark encompassing 5,000 instances across 25 tasks. Each task in the MTI BENCH involves 2 to 3 sub-tasks. As expected, we first demonstrate that MULTI-TASK INFERENCE reduces the total inference time by ×1.46 times in average since it does not require multiple inference calls. Interestingly, contrary to the expectation that LLMs would perform better when tasks are divided, we find that state-ofthe-art LLMs, such as LLAMA-2-CHAT-70B and GPT-4, show up to 7.3% and 12.4% improved performance with MULTI-TASK INFER-ENCE compared to SINGLE-TASK INFERENCE on the MTI BENCH. We release the MTI BENCH dataset and our code at this link 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40b771b4-4876-4f22-9e64-19e014350283Cited by top-tier papers5
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at OnceZhuoshi Pan, Qizhi Pei, Yu Li, Zinan Tang et al.ACL 2026 · 13 citations
- Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context LengthJingxuan Chen, Mohammad Taher Pilehvar, José Camacho-ColladosACL 2026 · 1 citation
- ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model CapabilitiesYifan Duan, Yihong Tang, Kehai Chen, Liqiang Nie et al.EMNLP 2025
- A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language ModelsRei Emura, Saku SugawaraACL 2026
- JailbreakLoRA: Your Downloaded LoRA from Sharing Platforms might be UnsafeFanjunduo Wei, Zhenheng Tang, Rongfei Zeng, Tongliang Liu et al.ICLR 2026
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu et al.ICLR 2025
- EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language ModelsTao Zou, Xinghua Zhang, Haiyang Yu, Minzheng Wang et al.EMNLP 2025 · 1 citation
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu et al.ACL 2025 · 5 citations
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu et al.ICLR 2025
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia et al.EMNLP 2024 · 2 citations
