GUIA uditor : Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices
Junlin Liu, Yifeng Cai, Shuai Wang, Zhineng Zhong, Shaofei Li, Jiacheng Liu, Yuanchun Li, Ziqi Zhang, Xiangqun Chen, Yao Guo, Ding Li
摘要
The proliferation of smart devices exposes children to online risks like grooming and financial scams that are deeply embedded within legitimate applications. Current approaches rely on automated prevention and detection, a paradigm that is fundamentally limited by its inherent fallibility. Whether rule-based or AI-driven, they inevitably produce false positives and negatives, failing to provide reliable protection. In this paper, we argue for a complementary, human-in-the-loop, post-hoc forensic paradigm. We present GUIAuditor, the first system designed to realize this vision by creating GUI Provenance: a queryable, semantic record of a child's interaction sequence. To generate this, GUIAuditor leverages a Multimodal Large Language Model (MLLM) to translate the temporal sequence of GUI events into a human-understandable narrative. To make this practical on mobile devices, a novel evidence distillation pipeline reduces the data requiring analysis by over 89.2% compared to periodic sampling approaches adopted by industry standards, with negligible impact on accuracy. On a new dataset of 295 interaction clips, GUIAuditor achieves a 95.23% Macro-F1 Score in logging significant events and, crucially, its two-stage forensic query engine successfully retrieves the correct evidence as the top result for over 90.20% of natural language questions. An end-to-end evaluation on three modern smartphones shows that the full pipeline, including on-device MLLM inference, adds 2.1W of power draw and 7.4s of per-event latency, with a peak memory footprint of ∼3.1GB. These results show that post-hoc GUI forensics can run on modern mobile devices and provide useful context for guardian-led safety review.
CCS Concepts: • Security and privacy → Usability in security and privacy; Mobile platform security; • Human-centered computing → Ubiquitous and mobile computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information FlowsSadegh Momeni Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar 等S&P 2019 · 被引用 550 次
- UI Dark Patterns and Where to Find Them: A Study on Mobile Applications and User PerceptionLinda Di Geronimo, Larissa Braz, Enrico Fregnan, Fabio Palomba 等CHI 2020 · 被引用 262 次
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong 等USENIX Security 2024 · 被引用 121 次
相关 Paper
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou 等CVPR 2025
- The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode ForensicsTimo SchefoldCCS 2026
- Towards Scalable and Interpretable Mobile App Risk Analysis via Large Language ModelsYu Yang, Zhenyuan Li, Xiandong Ran, Jiahao Liu 等ICSE 2026
- ForeDroid: Scenario-Aware Analysis for Android Malware Detection and ExplanationJiaming Li, Sen Chen, Chunlian Wu, Yuxin Zhang 等CCS 2025
- Screen after Previous Screens: Spatial-Temporal Recreation of Android App Displays from Memory ImagesBrendan Saltaformaggio, Rohit Bhatia, Xiangyu Zhang, Dongyan Xu 等USENIX Security 2016 · 被引用 32 次
