Semantic GUI Scene Learning and Video Alignment for Detecting Duplicate Video-based Bug Reports
Yanfu Yan, Nathan Cooper, Oscar Chaparro, Kevin Moran, Denys Poshyvanyk
摘要
Video-based bug reports are increasingly being used to document bugs for programs centered around a graphical user interface (GUI). However, developing automated techniques to manage video-based reports is challenging as it requires identifying and understanding often nuanced visual patterns that capture key information about a reported bug. In this paper, we aim to overcome these challenges by advancing the bug report management task of duplicate detection for video-based reports. To this end, we introduce a new approach, called Janus, that adapts the scene-learning capabilities of vision transformers to capture subtle visual and textual patterns that manifest on app UI screens --- which is key to differentiating between similar screens for accurate duplicate report detection. Janus also makes use of a video alignment technique capable of adaptive weighting of video frames to account for typical bug manifestation patterns. In a comprehensive evaluation on a benchmark containing 7,290 duplicate detection tasks derived from 270 video-based bug reports from 90 Android app bugs, the best configuration of our approach achieves an overall mRR/mAP of 89.8%/84.7%, and for the large majority of duplicate detection tasks, outperforms prior work by ≈9% to a statistically significant degree. Finally, we qualitatively illustrate how the scene-learning capabilities provided by Janus benefits its performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Toward the Automated Localization of Buggy Mobile App UIs from Bug DescriptionsAntu Saha, Yang Song, Junayed Mahmud, Ying Zhou 等ISSTA 2024 · 被引用 7 次
- GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific BugsRuofan Liu, Xiwen Teoh, Yun Lin, Guanjie Chen 等ISSTA 2025 · 被引用 5 次
- SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI AutomationDehai Zhao, Zhenchang Xing, Qinghua Lu, Xiwei Xu 等ICSE 2025 · 被引用 1 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- It Takes Two to TANGO: Combining Visual and Textual Information for Detecting Duplicate Video-Based Bug ReportsNathan Cooper, Carlos Bernal-Cárdenas, Oscar Chaparro, Kevin Moran 等ICSE 2021 · 被引用 2 次
- Translating video recordings of mobile app usages into replayable scenariosCarlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro 等ICSE 2020 · 被引用 61 次
- Read It, Don't Watch It: Captioning Bug Recordings AutomaticallySidong Feng, Mulong Xie, Yinxing Xue, Chunyang ChenICSE 2023 · 被引用 12 次
- ViBR: Automated Bug Replay from Video-Based Reports using Vision-Language ModelsSidong Feng, Dingbang Wang, Nikola Tomic, Tingting Yu 等FSE 2026 · 被引用 1 次
- On Using GUI Interaction Data to Improve Text Retrieval-based Bug LocalizationJunayed Mahmud, Nadeeshan De Silva, Safwat Ali Khan, Seyed Hooman Mostafavi 等ICSE 2024 · 被引用 12 次
