DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
Yiheng Bian, Yunpeng Song, Guiyu Ma, Rongrong Zhu, Zhongmin Cai
Abstract
Information seeking on mobile devices is often fragmented, trapping users in repetitive cycles of context switching and data re-entry, which increases cognitive load and disrupts workflow. Existing mobile agents provide limited cross-source integration and are largely opaque, presenting progress as a linear feed with few opportunities to intervene, steer, or take control. We present DroidRetriever, a transparent, steerable system for cross-source mobile information seeking. It accepts voice or typed input and the multi-LLM system decomposes the task, navigates to target pages, takes screenshots, and synthesizes a concise report with citation-linked screenshots. We make the process transparent through a progress dashboard combining sub-task progress and real-time exploration maps for seamless takeover. DroidRetriever also pauses on detected privacy or high-risk screens and prompts intervention. Across 35 tasks over 24 apps, experiments and user studies demonstrate improvements in coverage, transparency, and reduced workload. We release our code at https://github.com/AkimotoAyako/DroidRetriever.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on23
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-Of-The-Art and Future DirectionsMagdalena Wischnewski, Nicole C. Krämer, Emmanuel MüllerCHI 2023 · 135 citations
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
Related papers
- A Study of Cross-Session Cross-Device Search Within an Academic Digital LibrarySebastian Gomes, Miriam L. Boon, Orland HoeberSIGIR 2022 · 17 citations
- MindSearch: Mimicking Human Minds Elicits Deep AI SearcherZehui Chen, Kuikun Liu, Qiuchen Wang, Jiangning Liu et al.ICLR 2025 · 2 citations
- Curiosity Driven Knowledge Retrieval for Mobile AgentsSijia Li, Xiaoyu Tan, Shahir Ali, Niels Schmidt et al.WWW 2026 · 1 citation
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile DevicesQuanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du et al.ICCV 2025 · 7 citations
- Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile AutomationYuxiang Zhou, Jichang Li, Yanhao Zhang, Haonan Lu et al.AAAI 2026 · 1 citation
