Testing Deep Learning Libraries via Neurosymbolic Constraint Learning
M M Abid Naziri, Shinhae Kim, Feiran (Alex) Qin, Saikat Dutta, Marcelo d'Amorim
摘要
Deep Learning (DL) libraries (e.g., Pytorch) are popular in the development of AI applications. These libraries are complex and contain bugs. Researchers have proposed various bug-finding techniques for such libraries. Yet, there is much room for improvement.
A key challenge in testing DL libraries is the lack of API specifications. Prior testing approaches often inaccurately model the input specifications of DL APIs, resulting in missed valid inputs that could reveal bugs or false alarms due to invalid inputs.
To address this challenge, we develop Centaur-the first neurosymbolic technique to test DL library APIs using dynamically learned input constraints. Centaur leverages the key idea that formal API constraints can be learned from a small number of automatically generated seed inputs, and that the learned constraints can be solved using SMT solvers to generate valid and diverse test inputs to test the API.
We develop a novel grammar that represents first-order logic formulae over API parameters and expresses tensor-related properties (e.g., shape, tensor data types, etc.) as well as relational properties between parameters. We use the grammar to guide a Large Language Model (LLM) to enumerate syntactically correct candidate rules, which we then validate using the seed inputs. Further, we develop a custom refinement strategy to prune the set of learned rules to eliminate spurious or redundant rules. We use the learned constraints to systematically generate valid and diverse inputs for the API by integrating SMT-based solving with randomized sampling.
We evaluate Centaur for testing PyTorch and TensorFlow. Our results show that Centaur's constraints have a recall of 94.0% and a precision of 94.0% on average. In terms of coverage, Centaur covers 203, 150, and 9,608 more branches than TitanFuzz, ACETest and Pathfinder, respectively. Using Centaur, we also detect 26 new bugs in PyTorch and TensorFlow, 18 of which are confirmed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang 等ISSTA 2023 · 被引用 253 次
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu 等FSE 2020 · 被引用 165 次
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 被引用 91 次
- NNSmith: Generating Diverse and Valid Test Cases for Deep Learning CompilersJiawei Liu, Jinkun Lin, Fabian Ruffy, Cheng Tan 等ASPLOS 2023 · 被引用 90 次
相关 Paper
- ACETest: Automated Constraint Extraction for Testing Deep Learning OperatorsJingyi Shi, Yang Xiao, Yuekang Li, Yeting Li 等ISSTA 2023 · 被引用 24 次
- DocTer: documentation-guided fuzzing for testing deep learning API functionsDanning Xie, Yitong Li, Mijung Kim, Hung Viet Pham 等ISSTA 2022 · 被引用 72 次
- Fuzzing deep-learning libraries via automated relational API inferenceYinlin Deng, Chenyuan Yang, Anjiang Wei, Lingming ZhangFSE 2022 · 被引用 83 次
- Lightweight Concolic Testing via Path-Condition Synthesis for Deep Learning LibrariesSehoon Kim, Yonghyeon Kim, Dahyeon Park, Yuseok Jeon 等ICSE 2025 · 被引用 5 次
- Towards More Complete Constraints for Deep Learning Library Testing via Complementary Set Guided RefinementGwihwan Go, Chijin Zhou, Quan Zhang, Xiazijian Zou 等ISSTA 2024 · 被引用 2 次
