Co-Designing Traffic Control with NVMe-oF for Disaggregated Storage: A Comparative Study of Switched and Switchless SAN Architectures
Chendong Wang, Joontaek Oh, Ming Liu
摘要
Disaggregated storage is a pivotal component of today's cluster infrastructures. With the advent of high-bandwidth server interconnects and new NVMe form factors, commodity storage appliances are becoming denser, delivering tens of millions of IOPS. This calls for today's storage area network (SAN) fabric to expand the bandwidth capacity drastically. Industry practices tackle this issue via either (i) a scale-up approach, upgrading the per-port bandwidth in a switched SAN, or (ii) a scale-out strategy, integrating more paths in a switchless SAN. However, it is unclear which network architecture is more suitable for scaling storage disaggregation.
This paper presents a comparative study of switched and switchless SAN architectures from several angles. We begin by developing an experimental methodology that integrates both small-scale real-system prototypes and large-scale simulations, providing the flexibility needed to explore architectural trade-offs. We then characterize NVMe-oF I/O flows and co-design SAN traffic control mechanisms around these characteristics to improve I/O transmission efficiency in both settings. Our evaluation yields several key findings. First, the switchless SAN achieves throughput comparable to that of the switched SAN, despite involving additional routing hops, while simultaneously reducing latency through the use of multiple load-aware I/O paths that mitigate interference. Second, the switchless SAN reduces capital costs by obviating the need for expensive high-radix switches, scales effectively under heterogeneous I/O workloads, and avoids the single point of failure associated with top-of-rack (ToR) switches. Collectively, these results demonstrate that switchless SANs provide a compelling alternative to traditional switched designs for disaggregated storage environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Building A CSFQ-Inspired Transport for Switched CXL Memory PoolingZerui Guo, Emily Shriver, Ming LiuNSDI 2026 · 被引用 2 次
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
它引用的顶会 Paper23
- SP-PIFO: Approximating Push-In First-Out Behaviors using Strict-Priority QueuesAlbert Gran Alcoz, Alexander Dietmüller, Laurent VanbeverNSDI 2020 · 被引用 140 次
- Programmable Calendar Queues for High-speed Packet SchedulingNaveen Kr. Sharma, Chenxingyu Zhao, Ming Liu, Pravein G. Kannan 等NSDI 2020 · 被引用 119 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudRui Miao, Lingjun Zhu, Shu Ma, Kun Qian 等SIGCOMM 2022 · 被引用 79 次
- Rearchitecting Linux Storage Stack for µs Latency and High ThroughputJaehyun Hwang, Midhul Vuppalapati, Simon Peter, Rachit AgarwalOSDI 2021 · 被引用 63 次
相关 Paper
- Scalio: Scaling up DPU-based JBOF Key-value Store with NVMe-oF Target OffloadXun Sun, Mingxing Zhang, Yingdi Shan, Kang Chen 等OSDI 2025 · 被引用 2 次
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- NVMe-oAF: Towards Adaptive NVMe-oF for IO-Intensive Workloads on HPC CloudArjun Kashyap, Xiaoyi LuHPDC 2022 · 被引用 10 次
- CETOFS: A High-Performance File System with Host-Server Collaboration for Remote StorageWenqing Jia, Dejun Jiang, Jin XiongFAST 2026
- NeVerMore: Exploiting RDMA Mistakes in NVMe-oF Storage ApplicationsKonstantin Taranov, Benjamin Rothenberger, Daniele De Sensi, Adrian Perrig 等CCS 2022 · 被引用 8 次
