Natural Symbolic Execution-Based Testing for Big Data Analytics
Yaoxuan Wu, Ahmad Humayun, Muhammad Ali Gulzar, Miryung Kim
摘要
Symbolic execution is an automated test input generation technique that models individual program paths as logical constraints. However, the realism of concrete test inputs generated by SMT solvers often comes into question. Existing symbolic execution tools only seek arbitrary solutions for given path constraints. These constraints do not incorporate the naturalness of inputs that observe statistical distributions, range constraints, or preferred string constants. This results in unnatural-looking inputs that fail to emulate real-world data. In this paper, we extend symbolic execution with consideration for incorporating naturalness. Our key insight is that users typically understand the semantics of program inputs, such as the distribution of height or possible values of zipcode , which can be leveraged to advance the ability of symbolic execution to produce natural test inputs. We instantiate this idea in N atural S ym , a symbolic execution-based test generation tool for data-intensive scalable computing (DISC) applications. NaturalSym generates natural-looking data that mimics real-world distributions by utilizing user-provided input semantics to drastically enhance the naturalness of inputs, while preserving strong bug-finding potential. On DISC applications and commercial big data test benchmarks, N atural S ym achieves a higher degree of realism —as evidenced by a perplexity score 35.1 points lower on median, and detects 1.29× injected faults compared to the state-of-the-art symbolic executor for DISC, B ig T est . This is because B ig T est draws inputs purely based on the satisfiability of path constraints constructed from branch predicates, while N atural S ym is able to draw natural concrete values based on user-specified semantics and prioritize using these values in input generation. Our empirical results demonstrate that NaturalSym finds injected faults 47.8× more than N atural F uzz (a coverage-guided fuzzer) and 19.1× more than ChatGPT. Meanwhile, TestMiner (a mining-based approach) fails to detect any injected faults. N atural S ym is the first symbolic executor that combines the notion of input naturalness in symbolic path constraints during SMT-based input generation. We make our code available at https://github.com/UCLA-SEAL/NaturalSym .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 被引用 156 次
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney 等ICSE 2023 · 被引用 45 次
- CAT-LM Training Language Models on Aligned Code And TestsNikitha Rao, Kush Jain, Uri Alon, Claire Le Goues 等ASE 2023 · 被引用 37 次
相关 Paper
- NaturalFuzz: Natural Input Generation for Big Data AnalyticsAhmad Humayun, Yaoxuan Wu, Miryung Kim, Muhammad Ali GulzarASE 2023 · 被引用 2 次
- SYMTUNER: Maximizing the Power of Symbolic Execution by Adaptively Tuning External ParametersSooyoung Cha, Myungho Lee, Seokhyun Lee, Hakjoo OhICSE 2022 · 被引用 4 次
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye 等ASE 2020 · 被引用 27 次
- Generator Solving for Symbolic ExecutionSiwei Wei, Yan CaiICSE 2026
- Topseed: Learning Seed Selection Strategies for Symbolic Execution from ScratchJaehyeok Lee, Sooyoung ChaICSE 2025
