VideoAgentTrek: Computer-Use Pretraining from Unlabeled Videos
Dunjie Lu, Yiheng Xu, Junli Wang, Haoyuan Wu, Xinyuan Wang, Zekun Wang, Junlin Yang, Hongjin SU, Jixuan Chen, Junda Chen, Yuchen Mao, Junyang Lin
Abstract
Training computer-use agents requires massive amounts of GUI interaction data, but manually annotating action trajectories at scale is prohibitively expensive. We present VIDEOAGENTTREK, a scalable pipeline that automatically mines training data from publicly available screen-recorded videos at web scale, eliminating the need for manual annotation. Our approach addresses a key challenge: raw videos contain implicit demonstrations but lack explicit action labels. To solve this, we develop VIDEO2ACTION, an inverse dynamics module (IDM) with two components: (1) a video grounding model that detects and localizes GUI actions with precise temporal boundaries and context, and (2) an action-content recognizer that extracts structured parameters like click coordinates and typed text with high fidelity. Applied to 39,000 YouTube tutorial videos, our pipeline generates 1.52 million interaction steps automatically. We leverage this data through continued pretraining followed by supervised fine-tuning. On OSWorld-Verified, our approach improves task success rates from 9.3% (SFT-only baseline) to 15.8%, a 70% relative improvement. On AgentNetBench, step accuracy increases from 64.1% to 69.3%. Our results demonstrate that passive internet videos can be transformed into high-quality supervision for computer-use agents, providing a scalable alternative to expensive manual annotation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95ec98f6-4d22-4e17-baae-1c07515bce25Cited by top-tier papers1
Ask how each one uses itBuilds on11
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang et al.NeurIPS 2025 · 151 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- VideoGrounding-DINO: Towards Open-Vocabulary Spatio- Temporal Video GroundingSyed Talal Wasim, Muzammal Naseer, Salman H. Khan, Ming-Hsuan Yang et al.CVPR 2024 · 10 citations
Related papers
- Watch and Learn: Learning to Use Computers from Online VideosChan Hee Song, Yiwen Song, Palash Goyal, Yu Su et al.CVPR 2026 · 8 citations
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang et al.ICLR 2025
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent PretrainingWeimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue et al.ICML 2026 · 2 citations
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement LearningLongxi Gao, Li Zhang, Pengzhi Gao, Wei Liu et al.ICLR 2026 · 11 citations
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang et al.AAAI 2026 · 26 citations
