Turaco: Complexity-Guided Data Sampling for Training Neural Surrogates of Programs
Alex Renda, Yi Ding, Michael Carbin
摘要
Programmers and researchers are increasingly developing surrogates of programs, models of a subset of the observable behavior of a given program, to solve a variety of software development challenges. Programmers train surrogates from measurements of the behavior of a program on a dataset of input examples. A key challenge of surrogate construction is determining what training data to use to train a surrogate of a given program.
We present a methodology for sampling datasets to train neural-network-based surrogates of programs. We first characterize the proportion of data to sample from each region of a program's input space (corresponding to different execution paths of the program) based on the complexity of learning a surrogate of the corresponding execution path. We next provide a program analysis to determine the complexity of different paths in a program. We evaluate these results on a range of real-world programs, demonstrating that complexity-guided sampling results in empirical improvements in accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- NEUZZ: Efficient Fuzzing with Neural Program SmoothingDongdong She, Kexin Pei, Dave Epstein, Junfeng Yang 等S&P 2019 · 被引用 220 次
- DiffTune: Optimizing CPU Simulator Parameters with Learned Differentiable SurrogatesAlex Renda, Yishen Chen, Charith Mendis, Michael CarbinMICRO 2020 · 被引用 26 次
- One Network Fits All? Modular versus Monolithic Task Formulations in Neural NetworksAtish Agarwala, Abhimanyu Das, Brendan Juba, Rina Panigrahy 等ICLR 2021 · 被引用 3 次
相关 Paper
- Learning to Compile Programs to Neural NetworksLogan Weber, Jesse Michel, Alex Renda, Michael CarbinICML 2024 · 被引用 2 次
- Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code ExecutorsBohan Lyu, Siqiao Huang, Zichen Liang, Qian Sun 等EMNLP 2025
- Auto-HPCnet: An Automatic Framework to Build Neural Network-based Surrogate for High-Performance Computing ApplicationsWenqian Dong, Gokcen Kestor, Dong LiHPDC 2023 · 被引用 6 次
- HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisShraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick 等NeurIPS 2024 · 被引用 34 次
- Approximate Computing Through the Lens of Uncertainty QuantificationKonstantinos Parasyris, James Diffenderfer, Harshitha Menon, Ignacio Laguna 等SC 2022 · 被引用 5 次
