Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong, Seungone Kim
摘要
Large language models (LLMs) are typically prompted to follow a single instruction per inference call. In this work, we analyze whether LLMs also hold the capability to handle multiple instructions simultaneously, denoted as MULTI-TASK INFERENCE. For this purpose, we introduce the MTI BENCH (Multi-Task Inference Benchmark), a comprehensive evaluation benchmark encompassing 5,000 instances across 25 tasks. Each task in the MTI BENCH involves 2 to 3 sub-tasks. As expected, we first demonstrate that MULTI-TASK INFERENCE reduces the total inference time by ×1.46 times in average since it does not require multiple inference calls. Interestingly, contrary to the expectation that LLMs would perform better when tasks are divided, we find that state-ofthe-art LLMs, such as LLAMA-2-CHAT-70B and GPT-4, show up to 7.3% and 12.4% improved performance with MULTI-TASK INFER-ENCE compared to SINGLE-TASK INFERENCE on the MTI BENCH. We release the MTI BENCH dataset and our code at this link 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at OnceZhuoshi Pan, Qizhi Pei, Yu Li, Zinan Tang 等ACL 2026 · 被引用 13 次
- Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context LengthJingxuan Chen, Mohammad Taher Pilehvar, José Camacho-ColladosACL 2026 · 被引用 1 次
- ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model CapabilitiesYifan Duan, Yihong Tang, Kehai Chen, Liqiang Nie 等EMNLP 2025
- A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language ModelsRei Emura, Saku SugawaraACL 2026
- JailbreakLoRA: Your Downloaded LoRA from Sharing Platforms might be UnsafeFanjunduo Wei, Zhenheng Tang, Rongfei Zeng, Tongliang Liu 等ICLR 2026
它引用的顶会 Paper12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
相关 Paper
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu 等ICLR 2025
- EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language ModelsTao Zou, Xinghua Zhang, Haiyang Yu, Minzheng Wang 等EMNLP 2025 · 被引用 1 次
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu 等ACL 2025 · 被引用 5 次
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 等ICLR 2025
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia 等EMNLP 2024 · 被引用 2 次
