NetArena: Dynamic Benchmarks for AI Agents in Network Automation
Yajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi, Francis Y. Yan, Kevin Hsieh, Zaoxing Liu
摘要
As AI agents expand into high-stakes domains like network system operations, evaluating their real-world reliability becomes increasingly critical. However, existing benchmarks risk contamination due to static design, show high statistical variance from limited dataset size, and fail to reflect the complexity of production environments. We present NETARENA, a dynamic benchmark generation framework for network applications. NETARENA introduces a novel abstraction and unified interface that generalize across diverse tasks, enabling dynamic benchmarking despite the heterogeneity of network workloads. At runtime, users can generate unlimited queries on demand. NETARENA integrates with network emulators to measure correctness, safety, and latency during execution. We demonstrate NETARENA on three representative applications and find that (1) NETARENA significantly improves statistical reliability across AI agents, reducing confidence-interval overlap from 85% to 0, (2) agents achieve only 13-38% average performance (as low as 3%) for large-scale, realistic queries, and (3) it exposes more fine-grained behaviors that static, correctness-only benchmarks miss. NETARENA also enables use cases such as SFT and RL fine-tuning on network system tasks. Code is available at https://github.com/Froot-NetSys/NetArena .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak 等ICLR 2024 · 被引用 220 次
相关 Paper
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk EvaluationChangyi Li, Pengfei Lu, Xudong Pan, Fazl Barez 等ICML 2026 · 被引用 2 次
- ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksSaurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa 等ICML 2025
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 被引用 2 次
- GraphArena: Evaluating and Exploring Large Language Models on Graph ComputationJianheng Tang, Qifan Zhang, Yuhan Li, Nuo Chen 等ICLR 2025
