Scalable Parallel Flash Firmware for Many-core Architectures
Jie Zhang, Miryeong Kwon, Michael M. Swift, Myoungsoo Jung
Abstract
NVMe is designed to unshackle flash from a traditional storage bus by allowing hosts to employ many threads to achieve higher bandwidth. While NVMe enables users to fully exploit all levels of parallelism offered by modern SSDs, current firmware designs are not scalable and have difficulty in handling a large number of I/O requests in parallel due to its limited computation power and many hardware contentions.
We propose DeepFlash, a novel manycore-based storage platform that can process more than a million I/O requests in a second (1MIOPS) while hiding long latencies imposed by its internal flash media. Inspired by a parallel data analysis system, we design the firmware based on many-to-many threading model that can be scaled horizontally. The proposed DeepFlash can extract the maximum performance of the underlying flash memory complex by concurrently executing multiple firmware components across many cores within the device. To show its extreme parallel scalability, we implement DeepFlash on a many-core prototype processor that employs dozens of lightweight cores, analyze new challenges from parallel I/O processing and address the challenges by applying concurrency-aware optimizations. Our comprehensive evaluation reveals that DeepFlash can serve around 4.5 GB/s, while minimizing the CPU demand on microbenchmarks and real server workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bda8782d-8234-49d2-91c2-668a0677adfcCited by top-tier papers10
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- ZNS+: Advanced Zoned Namespace Interface for Supporting In-Storage Zone CompactionKyuhwa Han, Hyunho Gwak, Dongkun Shin, Jooyoung HwangOSDI 2021 · 107 citations
- Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNsShine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae et al.FAST 2021 · 44 citations
- Rearchitecting the TCP Stack for I/O-Offloaded Content DeliveryTaehyun Kim, Deondre Martin Ng, Junzhi Gong, Youngjin Kwon et al.NSDI 2023 · 42 citations
- Hardware/Software Co-Programmable Framework for Computational SSDs to Accelerate Deep Learning Service on Large-Scale GraphsMiryeong Kwon, Donghyun Gouk, Sangwon Lee, Myoungsoo JungFAST 2022 · 32 citations
Related papers
- PipeSSD: A Lock-free Pipelined SSD Firmware Design for Multi-core ArchitectureZelin Du, Shaoqi Li, Zixuan Huang, Jin Xue et al.DAC 2024
- What Modern NVMe Storage Can Do, And How To Exploit It: High-Performance I/O for High-Performance Storage EnginesGabriel Haas, Viktor LeisVLDB 2023 · 83 citations
- BypassD: Enabling fast userspace access to shared SSDsSujay Yadalam, Chloe Alverti, Vasileios Karakostas, Jayneel Gandhi et al.ASPLOS 2024 · 5 citations
- Daredevil: Rescue Your Flash Storage from Inflexible Kernel Storage StackJunzhe Li, Ran Shu, Jiayi Lin, Qingyu Zhang et al.EuroSys 2025 · 2 citations
- Optimizing Memory-mapped I/O for Fast Storage DevicesAnastasios Papagiannis, Giorgos Xanthakis, Giorgos Saloustros, Manolis Marazakis et al.USENIX ATC 2020 · 68 citations
