Beware the evolving 'intelligent' web service! an integration architecture tactic to guard AI-first components
Alex Cummaudo, Scott Barnett, Rajesh Vasa, John C. Grundy, Mohamed Abdelrazek
Abstract
Intelligent services provide the power of AI to developers via simple RESTful API endpoints, abstracting away many complexities of machine learning. However, most of these intelligent services---such as computer vision---continually learn with time. When the internals within the abstracted 'black box' become hidden and evolve, pitfalls emerge in the robustness of applications that depend on these evolving services. Without adapting the way developers plan and construct projects reliant on intelligent services, significant gaps and risks result in both project planning and development. Therefore, how can software engineers best mitigate software evolution risk moving forward, thereby ensuring that their own applications maintain quality? Our proposal is an architectural tactic designed to improve intelligent service-dependent software robustness. The tactic involves creating an application-specific benchmark dataset baselined against an intelligent service, enabling evolutionary behaviour changes to be mitigated. A technical evaluation of our implementation of this architecture demonstrates how the tactic can identify 1,054 cases of substantial confidence evolution and 2,461 cases of substantial changes to response label sets using a dataset consisting of 331 images that evolve when sent to a service.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaa7f6c6-b273-4a37-a01f-d901ef3001b6Cited by top-tier papers1
Ask how each one uses itBuilds on2
- With Great Training Comes Great Vulnerability: Practical Attacks against Transfer LearningBolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng et al.USENIX Security 2018 · 126 citations
- Interpreting cloud computer vision pain-points: a mining study of stack overflowAlex Cummaudo, Rajesh Vasa, Scott Barnett, John C. Grundy et al.ICSE 2020 · 26 citations
Related papers
- On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)Weichang Liu, Junwei Zhang, Yuqing Niu, Bo ZhouISSTA 2026
- PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI AgentsQihang Cen, Tianshuo Cong, Da Song, Xinlei He et al.CCS 2026
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 193 citations
- Evolution of Benchmark: Black-Box Optimization Benchmark Design through Large Language ModelChen Wang, Sijie Ma, Zeyuan Ma, Yue-Jiao GongICML 2026
- GUIDER: GUI structure and vision co-guided test script repair for Android appsTongtong Xu, Minxue Pan, Yu Pei, Guiyin Li et al.ISSTA 2021 · 30 citations
