PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Apurva Gandhi, Vishwas Suryanarayanan, Raja Anwar, Firoz Shaik, Shubhang Desai, Thong Nguyen, Muhammad Raza, Vishal Chowdhary, Graham Neubig
摘要
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft Pow-erPoint is among the most widely adopted and feature-rich environments for presentation creation. We introduce PPT-EVAL, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty. A central challenge in this domain is evaluation: tasks are complex, multimodal, and often admit many valid solutions. Moreover, today's agents frequently make only partial progress, which binary success metrics fail to capture. To address this, we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks, taking inspiration from and building on past works for rubric-based evaluation. These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's τ b correlation of 0.77 with human judgments. We find that existing frontier agents still struggle with solving PowerPoint tasks, with strong models like Claude-4.5-Opus achieving only a 45% success rate and an average partial score of 57%. The benchmark repository is located at: https: //microsoft.github.io/ppteval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang 等NeurIPS 2025 · 被引用 151 次
- Checklists Are Better Than Reward Models For Aligning Language ModelsVijay Viswanathan, Yanchao Sun, Xiang Kong, Meng Cao 等NeurIPS 2025 · 被引用 127 次
- Go-Browse: Training Web Agents with Structured ExplorationApurva Gandhi, Graham NeubigICLR 2026 · 被引用 30 次
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang 等MobiCom 2025 · 被引用 6 次
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleRogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont 等ICML 2025
相关 Paper
- PPTAgent: Generating and Evaluating Presentations Beyond Text-to-SlidesHao Zheng, Xinyan Guan, Hao Kong, Wenkai Zhang 等EMNLP 2025 · 被引用 3 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao 等ICLR 2026 · 被引用 30 次
- PSBench: Editing Image via GUI Agents in PhotoshopYinuo Zhang, Zian Cheng, Ziya Zhao, Zongyu Li 等ICML 2026
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
