TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
Yikai Zhang, Siyu Yuan, Caiyu Hu, Kyle Richardson, Yanghua Xiao, Jiangjie Chen
Abstract
Despite remarkable advancements in emulating human-like behavior through Large Language Models (LLMs), current textual simulations do not adequately address the notion of time. To this end, we introduce TIMEARENA, a novel textual simulated environment that incorporates complex temporal dynamics and constraints that better reflect real-life planning scenarios. In TIMEARENA, agents are asked to complete multiple tasks as soon as possible, allowing for parallel processing to save time. We implement the dependency between actions, the time duration for each action, and the occupancy of the agent and the objects in the environment. TIMEARENA grounds to 30 real-world tasks in cooking, household activity, and laboratory work. We conduct extensive experiments with various LLMs using TIMEARENA. Our findings reveal that even the most powerful models, e.g., GPT-4, still lag behind humans in effective multitasking, underscoring the need for enhanced temporal awareness in the development of language agents. pendencies, requiring agents to strategize and prior-042 itize based on time constraints and task completion 043 progress. 2) Agent Occupancy: Agents will be oc-044 cupied by certain actions thus they might be unable 045 to perform other actions at the same time. 3) Ob-046 ject Occupancy: Some objects might be occupied 047 for some time, and agents must use available ob-048 jects in the environment for the tasks. These factors 049 are common in real-life but are seldom addressed 050 by current textual simulations. 051 To help illustrate, Figure 1 shows an example 052 of completing make tea (Task 1) and wash clothes 053 (Task 2). The actions of each task might depend on 054 previous actions, e.g., agents must boil water be-055 fore make tea, and each action takes a duration in 056 time, e.g., wash cup takes 5 minutes. In particular, 057 2 Related Work 110 Simulation-based Evaluation For language 111 Agents With the great success of LLMs (Ope-112 nAI, 2022, 2023; Team and Google, 2023), recent 113 works have shifted the focus from traditional NLP 114 tasks to explore language agents in simulation en-115 vironments that mimic real-world scenarios (Wu 116
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu et al.ACL 2026 · 21 citations
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long et al.ICLR 2026 · 8 citations
Builds on7
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang et al.ICLR 2024 · 594 citations
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 24 citations
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu et al.ACL 2024 · 12 citations
- SituatedQA: Incorporating Extra-Linguistic Contexts into QAMichael J. Q. Zhang, Eunsol ChoiEMNLP 2021 · 2 citations
Related papers
- PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool LearningZihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan et al.ACL 2025 · 3 citations
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont et al.ICML 2025
- How Well Can LLMs Negotiate? NegotiationArena Platform and AnalysisFederico Bianchi, Patrick John Chia, Mert Yüksekgönül, Jacopo Tagliabue et al.ICML 2024 · 90 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
