TraceSplitter: a new paradigm for downscaling traces
Sultan Mahmud Sajal, Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen
摘要
Realistic experimentation is a key component of systems research and industry prototyping, but experimental clusters are often too small to replay the high traffic rates found in production traces. Thus, it is often necessary to downscale traces to lower their arrival rate, and researchers/practitioners generally do this in an ad-hoc manner. For example, one practice is to multiply all arrival timestamps in a trace by a scaling factor to spread the load across a longer timespan. However, temporal patterns are skewed by this approach, which may lead to inappropriate conclusions about some system properties (e.g., the agility of auto-scaling). Another popular approach is to count the number of arrivals in fixed-sized time intervals and scale it according to some modeling assumptions. However, such approaches can eliminate or exaggerate the fine-grained burstiness in the trace depending on the time interval length.
The goal of this paper is to demonstrate the drawbacks of common downscaling techniques and propose new methods for realistically downscaling traces. We introduce a new paradigm for scaling traces that splits an original trace into multiple downscaled traces to accurately capture the characteristics of the original trace. Our key insight is that production traces are often generated by a cluster of service instances sitting behind a load balancer; by mimicking the load balancing used to split load across these instances, we can similarly split the production trace in a manner that captures the workload experienced by each service instance. Using production traces, synthetic traces, and a case study of an auto-scaling system, we identify and evaluate a variety of scenarios that show how our approach is superior to current approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud ProviderJiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen 等USENIX ATC 2025 · 被引用 70 次
- TraceUpscaler: Upscaling Traces to Evaluate Systems at High LoadSultan Mahmud Sajal, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha SenEuroSys 2024 · 被引用 4 次
- Avoiding the Ordering Trap in Systems Performance MeasurementDmitry Duplyakin, Nikhil Ramesh, Carina Imburgia, Hamza Fathallah Al Sheikh 等USENIX ATC 2023 · 被引用 3 次
它引用的顶会 Paper3
- Borg: the next generationMuhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque 等EuroSys 2020 · 被引用 323 次
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou 等EuroSys 2020 · 被引用 49 次
- Waiting game: optimally provisioning fixed resources for cloud-enabled schedulersPradeep Ambati, Noman Bashir, David Irwin, Prashant J. ShenoySC 2020 · 被引用 14 次
相关 Paper
- FaaSRail: Employing Real Workloads to Generate Representative Load for Serverless ResearchChristos Katsakioris, Chloe Alverti, Konstantinos Nikas, Dimitrios Siakavaras 等HPDC 2024 · 被引用 5 次
- PASS: Predictive Auto-Scaling System for Large-scale Enterprise Web ApplicationsYunda Guo, Jiake Ge, Panfeng Guo, Yunpeng Chai 等WWW 2024 · 被引用 10 次
- Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability AnalysisHaiyu Huang, Cheng Chen, Kunyi Chen, Pengfei Chen 等ASPLOS 2025 · 被引用 6 次
- Practical Analysis of Replication-Based SystemsFlorin Ciucu, Felix Poloczek, Lydia Y. Chen, Martin ChanINFOCOM 2021 · 被引用 6 次
- Robust Auto-Scaling with Probabilistic Workload Forecasting for Cloud DatabasesHaitian Hang, Xiu Tang, Jianling Sun, Lingfeng Bao 等ICDE 2024 · 被引用 11 次
