QEI: Query Acceleration Can be Generic and Efficient in the Cloud
Yifan Yuan, Yipeng Wang, Ren Wang, Rangeen Basu Roy Chowdhury, Charlie Tai, Nam Sung Kim
Abstract
Data query operations of different data structures are ubiquitous and critical in today's data center infrastructures and applications. However, query operations are not always performance-optimal to be executed on general-purpose CPU cores. These operations exhibit insufficient memory-level parallelism and frontend bottlenecks due to unstructured control flow. Furthermore, the data access patterns are not cache- or prefetch-friendly. Based on our performance analysis on a commodity server, query operations can consume a large percentage of the CPU cycles in various modern cloud workloads. Existing accelerator solutions for query operations do not strike a balance between their generality, scalability, latency, and hardware complexity. In this paper, we propose QEI, a generic, integrated, and efficient acceleration solution for various data structure queries. We first abstract the query operations to a few regular steps and map them to a simple and hardware-friendly configurable finite automaton model. Based on this model, we develop the QEI architecture that allows multiple query operations to execute in parallel to maximize throughput. We also propose a novel way to integrate the accelerator into the CPU that balances performance, latency, and hardware cost. QEI keeps the main control logic near the L2 cache to leverage existing hardware resources in the core while distributing the data-intensive comparison logic to each last-level cache slice for higher parallelism. Our results with five representative data center workloads show that QEI can achieve 6.5× 11.2× performance improvement in various scenarios with low overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bacf14a0-3044-424c-95cd-4843dfc6f92eCited by top-tier papers1
Ask how each one uses itRelated papers
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng et al.SC 2022 · 7 citations
- FineStream: Fine-Grained Window-Based Stream Processing on CPU-GPU Integrated ArchitecturesFeng Zhang, Lin Yang, Shuhao Zhang, Bingsheng He et al.USENIX ATC 2020 · 41 citations
- Accelerating database analytic query workloads using an associative processorHelena Caminal, Yannis Chronis, Tianshu Wu, Jignesh M. Patel et al.ISCA 2022 · 19 citations
- RayDB: Building Databases with Ray Tracing CoresXuri Shi, Kai Zhang, X. Sean Wang, Xiaodong Zhang et al.VLDB 2026 · 3 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
