EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
Yufei He, Juncheng Liu, Yue Liu, Yibo Li, Tri Cao, Zhiyuan Hu, Xinxing Xu, Bryan Hooi
摘要
A fundamental limitation of current AI agents is their inability to learn complex skills on the fly at test time, often behaving like "clever but clueless interns" in novel environments. This severely limits their practical utility. To systematically measure and drive progress on this challenge, we first introduce the Jericho Test-Time Learning (J-TTL) benchmark. J-TTL is a new evaluation setup where an agent must play the same game for several consecutive episodes, attempting to improve its performance from one episode to the next. On J-TTL, we find that existing adaptation methods like reflection, memory, or reinforcement learning struggle. To address the challenges posed by our benchmark, we present EvoTest 1 , an evolutionary test-time learning framework that improves an agent without any fine-tuning or gradients-by evolving the entire agentic system after every episode. EvoTest has two roles: the Actor Agent, which plays the game, and the Evolver Agent, which analyzes the episode transcript to propose a revised configuration for the next run. This configuration rewrites the prompt, updates memory by logging effective state-action choices, tunes hyperparameters, and learns the tool-use routines. On our J-TTL benchmark, EvoTest consistently increases performance, outperforming not only reflection and memory-only baselines but also more complex online fine-tuning methods. Notably, our method is the only one capable of winning two games (Detective and Library), while all baselines fail to win any.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Promptbreeder: Self-Referential Self-Improvement via Prompt EvolutionChrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero 等ICML 2024 · 被引用 432 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
相关 Paper
- Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-ImprovementXinmeng Hou, Bohao Qu, Wuqi Wang, Peiliang Gong 等ACL 2026 · 被引用 1 次
- MemEvolve: Meta-Evolution of Agent Memory SystemsGuibin Zhang, Haotian Ren, Chong Zhan, Junhao Wang 等ICML 2026 · 被引用 69 次
- EscapeBench: Towards Advancing Creative Intelligence of Language Model AgentsCheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He 等ACL 2025
- CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative TournamentsLingyue Fu, Xin Ding, Linyue Pan, Yaoming Zhu 等ICML 2026 · 被引用 3 次
- Timely Machine: Awareness of Time Makes Test-Time Scaling AgenticYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li 等ACL 2026 · 被引用 3 次
