TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
Yikai Zhang, Siyu Yuan, Caiyu Hu, Kyle Richardson, Yanghua Xiao, Jiangjie Chen
摘要
Despite remarkable advancements in emulating human-like behavior through Large Language Models (LLMs), current textual simulations do not adequately address the notion of time. To this end, we introduce TIMEARENA, a novel textual simulated environment that incorporates complex temporal dynamics and constraints that better reflect real-life planning scenarios. In TIMEARENA, agents are asked to complete multiple tasks as soon as possible, allowing for parallel processing to save time. We implement the dependency between actions, the time duration for each action, and the occupancy of the agent and the objects in the environment. TIMEARENA grounds to 30 real-world tasks in cooking, household activity, and laboratory work. We conduct extensive experiments with various LLMs using TIMEARENA. Our findings reveal that even the most powerful models, e.g., GPT-4, still lag behind humans in effective multitasking, underscoring the need for enhanced temporal awareness in the development of language agents. pendencies, requiring agents to strategize and prior-042 itize based on time constraints and task completion 043 progress. 2) Agent Occupancy: Agents will be oc-044 cupied by certain actions thus they might be unable 045 to perform other actions at the same time. 3) Ob-046 ject Occupancy: Some objects might be occupied 047 for some time, and agents must use available ob-048 jects in the environment for the tasks. These factors 049 are common in real-life but are seldom addressed 050 by current textual simulations. 051 To help illustrate, Figure 1 shows an example 052 of completing make tea (Task 1) and wash clothes 053 (Task 2). The actions of each task might depend on 054 previous actions, e.g., agents must boil water be-055 fore make tea, and each action takes a duration in 056 time, e.g., wash cup takes 5 minutes. In particular, 057 2 Related Work 110 Simulation-based Evaluation For language 111 Agents With the great success of LLMs (Ope-112 nAI, 2022, 2023; Team and Google, 2023), recent 113 works have shifted the focus from traditional NLP 114 tasks to explore language agents in simulation en-115 vironments that mimic real-world scenarios (Wu 116
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu 等ACL 2026 · 被引用 21 次
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long 等ICLR 2026 · 被引用 8 次
它引用的顶会 Paper7
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 等ICLR 2024 · 被引用 594 次
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 被引用 24 次
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu 等ACL 2024 · 被引用 12 次
- SituatedQA: Incorporating Extra-Linguistic Contexts into QAMichael J. Q. Zhang, Eunsol ChoiEMNLP 2021 · 被引用 2 次
相关 Paper
- PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool LearningZihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan 等ACL 2025 · 被引用 3 次
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont 等ICML 2025
- How Well Can LLMs Negotiate? NegotiationArena Platform and AnalysisFederico Bianchi, Patrick John Chia, Mert Yüksekgönül, Jacopo Tagliabue 等ICML 2024 · 被引用 90 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur 等ACL 2024 · 被引用 25 次
