It Takes Two to TANGO: Combining Visual and Textual Information for Detecting Duplicate Video-Based Bug Reports
Nathan Cooper, Carlos Bernal-Cárdenas, Oscar Chaparro, Kevin Moran, Denys Poshyvanyk
Abstract
When a bug manifests in a user-facing application, it is likely to be exposed through the graphical user interface (GUI). Given the importance of visual information to the process of identifying and understanding such bugs, users are increasingly making use of screenshots and screen-recordings as a means to report issues to developers. However, when such information is reported en masse, such as during crowd-sourced testing, managing these artifacts can be a time-consuming process. As the reporting of screen-recordings in particular becomes more popular, developers are likely to face challenges related to manually identifying videos that depict duplicate bugs. Due to their graphical nature, screen-recordings present challenges for automated analysis that preclude the use of current duplicate bug report detection techniques. To overcome these challenges and aid developers in this task, this paper presents Tango, a duplicate detection technique that operates purely on video-based bug reports by leveraging both visual and textual information. Tango combines tailored computer vision techniques, optical character recognition, and text retrieval. We evaluated multiple configurations of Tango in a comprehensive empirical evaluation on 4,860 duplicate detection tasks that involved a total of 180 screen-recordings from six Android apps. Additionally, we conducted a user study investigating the effort required for developers to manually detect duplicate video-based bug reports and compared this to the effort required to use Tango. The results reveal that Tango's optimal configuration is highly effective at detecting duplicate video-based bug reports, accurately ranking target duplicate videos in the top-2 returned results in 83% of the tasks. Additionally, our user study shows that, on average, Tango can reduce developer effort by over 60%, illustrating its practicality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2daae2ca-8e53-49b5-90a7-08cb9ea60b2cCited by top-tier papers14
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng et al.CHI 2023 · 49 citations
- GIFdroid: Automated Replay of Visual Bug Reports for Android AppsSidong Feng, Chunyang ChenICSE 2022 · 38 citations
- Toward interactive bug reporting for (android app) end-usersYang Song, Junayed Mahmud, Ying Zhou, Oscar Chaparro et al.FSE 2022 · 29 citations
- Where is Your App Frustrating Users?Yawen Wang, Junjie Wang, Hongyu Zhang, Xuran Ming et al.ICSE 2022 · 23 citations
- Never-ending Learning of User InterfacesJason Wu, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin et al.UIST 2023 · 17 citations
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- Translating video recordings of mobile app usages into replayable scenariosCarlos Bernal-Cárdenas, Nathan Cooper, Kevin Moran, Oscar Chaparro et al.ICSE 2020 · 61 citations
- Seenomaly: vision-based linting of GUI animation effects against design-don't guidelinesDehai Zhao, Zhenchang Xing, Chunyang Chen, Xiwei Xu et al.ICSE 2020 · 55 citations
Related papers
- Semantic GUI Scene Learning and Video Alignment for Detecting Duplicate Video-based Bug ReportsYanfu Yan, Nathan Cooper, Oscar Chaparro, Kevin Moran et al.ICSE 2024 · 7 citations
- On Using GUI Interaction Data to Improve Text Retrieval-based Bug LocalizationJunayed Mahmud, Nadeeshan De Silva, Safwat Ali Khan, Seyed Hooman Mostafavi et al.ICSE 2024 · 12 citations
- ViBR: Automated Bug Replay from Video-Based Reports using Vision-Language ModelsSidong Feng, Dingbang Wang, Nikola Tomic, Tingting Yu et al.FSE 2026 · 1 citation
- Read It, Don't Watch It: Captioning Bug Recordings AutomaticallySidong Feng, Mulong Xie, Yinxing Xue, Chunyang ChenICSE 2023 · 12 citations
- Owl Eyes: Spotting UI Display Issues via Visual UnderstandingZhe Liu, Chunyang Chen, Junjie Wang, Yuekai Huang et al.ASE 2020 · 79 citations
