TorchQL: A Programming Framework for Integrity Constraints in Machine Learning
Aaditya Naik, Adam Stein, Yinjun Wu, Mayur Naik, Eric Wong
摘要
Finding errors in machine learning applications requires a thorough exploration of their behavior over data. Existing approaches used by practitioners are often ad-hoc and lack the abstractions needed to scale this process. We present TorchQL, a programming framework to evaluate and improve the correctness of machine learning applications. TorchQL allows users to write queries to specify and check integrity constraints over machine learning models and datasets. It seamlessly integrates relational algebra with functional programming to allow for highly expressive queries using only eight intuitive operators. We evaluate TorchQL on diverse use-cases including finding critical temporal inconsistencies in objects detected across video frames in autonomous driving, finding data imputation errors in time-series medical records, finding data labeling errors in realworld images, and evaluating biases and constraining outputs of language models. Our experiments show that TorchQL enables up to 13x faster query executions than baselines like Pandas and MongoDB, and up to 40% shorter queries than native Python. We also conduct a user study and find that TorchQL is natural enough for developers familiar with Python to specify complex integrity constraints.
CCS Concepts: • Software and its engineering → Domain specific languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DISCRET: Synthesizing Faithful Explanations For Treatment Effect EstimationYinjun Wu, Mayank Keoliya, Kan Chen, Neelay Velingker 等ICML 2024 · 被引用 3 次
- DOLPHIN: A Programmable Framework for Scalable Neurosymbolic LearningAaditya Naik, Jason Liu, Claire Wang, Amish Sethi 等ICML 2025
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 被引用 1,245 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- Generative Models for Effective ML on Private, Decentralized DatasetsSean Augenstein, H. Brendan McMahan, Daniel Ramage, Swaroop Ramaswamy 等ICLR 2020 · 被引用 207 次
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 被引用 144 次
相关 Paper
- Guardrail: Automated Integrity Constraint Synthesis From Noisy DataPingchuan Ma, Zhaoyu Wang, Zhenlan Ji, Zongjie Li 等SIGMOD 2026 · 被引用 1 次
- ExAIS: Executable AI SemanticsRichard Schumi, Jun SunICSE 2022 · 被引用 5 次
- Can LLMs Implicitly Learn Numeric Parameter Constraints in Data Science APIs?Yinlin Deng, Chunqiu Steven Xia, Zhezhen Cao, Meiziniu Li 等NeurIPS 2024 · 被引用 4 次
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQLYue Gong, Chuan Lei, Xiao Qin, Kapil Vaidya 等NeurIPS 2025 · 被引用 21 次
- DocTer: documentation-guided fuzzing for testing deep learning API functionsDanning Xie, Yitong Li, Mijung Kim, Hung Viet Pham 等ISSTA 2022 · 被引用 72 次
