Nüwa: A Generative Control Plane for AI Network Simulation
Wenkai Li, Ran Shu, Peng Zhang, Yiren Zhao, Danfeng Shan, Yongqiang Xiong
Abstract
Network simulation plays a critical role in improving the efficiency of large-scale AI clusters for design validation, parameter tuning, and protocol development. However, high-fidelity network simulation becomes prohibitively slow at scale, especially when running large batches of experiments on topologies with tens or hundreds of thousands of accelerators. We observe that a key bottleneck comes from the control plane. Existing network simulators typically compute routes and install forwarding tables at initialization, which can consume hundreds of GB of memory before packet-event execution begins and limit overall simulation throughput. In this paper, we present Nüwa, which views routing as a compilation problem, it leverages the hierarchical and symmetric structure common in AI fabrics and compiles a declarative topology description together with routing policies into compact forwarding artifacts that are fast to generate and efficient to look up. Evaluations show that Nüwa can reduce simulation initialization time from hours to only 25 seconds for a 65,536-GPU cluster. For end-to-end simulation time, Nüwa takes only 20% of that required by existing approaches in a 40K+ GPU cluster, and Nüwa can scale to a 221,184-GPU cluster.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 42fc0faa-e793-4dac-b7a3-be5fd245b3cfRelated papers
- Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at ScaleZhaodong Wang, Satyajeet Singh Ahuja, Xu Zhang, Max Noormohammadpour et al.NSDI 2026 · 3 citations
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
- Zettafly: A Network Topology with Flexible Non-blocking Regions for Large-scale AI and HPC SystemsDezun Dong, Ziyu Wang, Fei LeiISCA 2025 · 7 citations
- TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine LearningWilliam Won, Midhilesh Elavazhagan, Sudarshan Srinivasan, Swati Gupta et al.MICRO 2024 · 36 citations
- Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel SimulationFei Gui, Kaihui Gao, Li Chen, Dan Li et al.NSDI 2025 · 27 citations
