Finding data compatibility bugs with JSON subschema checking
Andrew Habib, Avraham Shinnar, Martin Hirzel, Michael Pradel
摘要
JSON is a data format used pervasively in web APIs, cloud computing, NoSQL databases, and increasingly also machine learning. To ensure that JSON data is compatible with an application, one can define a JSON schema and use a validator to check data against the schema. However, because validation can happen only once concrete data occurs during an execution, it may detect data compatibility bugs too late or not at all. Examples include evolving the schema for a web API, which may unexpectedly break client applications, or accidentally running a machine learning pipeline on incorrect data. This paper presents a novel way of detecting a class of data compatibility bugs via JSON subschema checking. Subschema checks find bugs before concrete JSON data is available and across all possible data specified by a schema. For example, one can check if evolving a schema would break API clients or if two components of a machine learning pipeline have incompatible expectations about data. Deciding whether one JSON schema is a subschema of another is non-trivial because the JSON Schema specification language is rich. Our key insight to address this challenge is to first reduce the richness of schemas by canonicalizing and simplifying them, and to then reason about the subschema question on simpler schema fragments using type-specific checkers. We apply our subschema checker to thousands of real-world schemas from different domains. In all experiments, the approach is correct whenever it gives an answer (100% precision and correctness), which is the case for most schema pairs (93.5% recall), clearly outperforming the state-of-the-art tool. Moreover, the approach reveals 43 previously unknown bugs in popular software, most of which have already been fixed, showing that JSON subschema checking helps finding data compatibility bugs early.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Pipeline Combinators for Gradual AutoMLGuillaume Baudart, Martin Hirzel, Kiran Kate, Parikshit Ram 等NeurIPS 2021 · 被引用 27 次
- Witness Generation for JSON SchemaLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli 等VLDB 2022 · 被引用 14 次
- Exploiting Structure in Regular Expression QueriesLing Zhang, Shaleen Deep, Avrilia Floratou, Anja Gruenheid 等SIGMOD 2023 · 被引用 3 次
- Streaming Validation of JSON Documents Against SchemasAlexis Le Glaunec, Angela W. Li, Konstantinos MamourasVLDB 2026 · 被引用 2 次
- The raise of machine learning hyperparameter constraints in Python codeIngkarat Rak-amnouykit, Ana L. Milanova, Guillaume Baudart, Martin Hirzel 等ISSTA 2022 · 被引用 1 次
相关 Paper
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt 等SIGMOD 2021 · 被引用 18 次
- Blaze: Compiling JSON Schema for 10x Faster ValidationMichael Mior, Juan Cruz ViottiVLDB 2026 · 被引用 1 次
- Validation of Modern JSON Schema: Formalization and ComplexityLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli 等POPL 2024 · 被引用 19 次
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 被引用 25 次
- Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database ConstructionJiansen Song, Wensheng Dou, Yingying Zheng, Yu Gao 等VLDB 2026
