Finding data compatibility bugs with JSON subschema checking
Andrew Habib, Avraham Shinnar, Martin Hirzel, Michael Pradel
Abstract
JSON is a data format used pervasively in web APIs, cloud computing, NoSQL databases, and increasingly also machine learning. To ensure that JSON data is compatible with an application, one can define a JSON schema and use a validator to check data against the schema. However, because validation can happen only once concrete data occurs during an execution, it may detect data compatibility bugs too late or not at all. Examples include evolving the schema for a web API, which may unexpectedly break client applications, or accidentally running a machine learning pipeline on incorrect data. This paper presents a novel way of detecting a class of data compatibility bugs via JSON subschema checking. Subschema checks find bugs before concrete JSON data is available and across all possible data specified by a schema. For example, one can check if evolving a schema would break API clients or if two components of a machine learning pipeline have incompatible expectations about data. Deciding whether one JSON schema is a subschema of another is non-trivial because the JSON Schema specification language is rich. Our key insight to address this challenge is to first reduce the richness of schemas by canonicalizing and simplifying them, and to then reason about the subschema question on simpler schema fragments using type-specific checkers. We apply our subschema checker to thousands of real-world schemas from different domains. In all experiments, the approach is correct whenever it gives an answer (100% precision and correctness), which is the case for most schema pairs (93.5% recall), clearly outperforming the state-of-the-art tool. Moreover, the approach reveals 43 previously unknown bugs in popular software, most of which have already been fixed, showing that JSON subschema checking helps finding data compatibility bugs early.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2b7edc6-9a43-4f1a-bc78-258ac036a5daCited by top-tier papers6
- Pipeline Combinators for Gradual AutoMLGuillaume Baudart, Martin Hirzel, Kiran Kate, Parikshit Ram et al.NeurIPS 2021 · 27 citations
- Witness Generation for JSON SchemaLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli et al.VLDB 2022 · 14 citations
- Exploiting Structure in Regular Expression QueriesLing Zhang, Shaleen Deep, Avrilia Floratou, Anja Gruenheid et al.SIGMOD 2023 · 3 citations
- Streaming Validation of JSON Documents Against SchemasAlexis Le Glaunec, Angela W. Li, Konstantinos MamourasVLDB 2026 · 2 citations
- The raise of machine learning hyperparameter constraints in Python codeIngkarat Rak-amnouykit, Ana L. Milanova, Guillaume Baudart, Martin Hirzel et al.ISSTA 2022 · 1 citation
Related papers
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt et al.SIGMOD 2021 · 18 citations
- Blaze: Compiling JSON Schema for 10x Faster ValidationMichael Mior, Juan Cruz ViottiVLDB 2026 · 1 citation
- Validation of Modern JSON Schema: Formalization and ComplexityLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli et al.POPL 2024 · 19 citations
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 25 citations
- Detecting Data-Type-Related Logic Bugs in Relational DBMSs via Compatible Database ConstructionJiansen Song, Wensheng Dou, Yingying Zheng, Yu Gao et al.VLDB 2026
