FormAct: Agentic Source Editing for Rich-Format Document Generation
Eugene Yu, Xingxing Zhang, Yuan Xia, Tao Ge, XWang, FNU Kartik, Vishwas Suryanarayanan, Cheng Yang, Amanda Jiang, Jiayu Ding, Xiangyu Wong, Tengchao Lv
摘要
Rich-format documents are essential for everyday operations yet costly to author, motivating the need for automated generation to enhance productivity. To this end, we present FormAct, an agentic system that generates professional rich-format documents from scratch. FormAct operates on an HTML source representation and performs iterative source refinement with an editing agent that invokes a suite of tools, including a syntax-aware source editor and a template retriever, and a review agent that critiques rendered pages to guide refinement. Additionally, we incorporate edit-triggered context compression to maintain a bounded working context and keep multi-round editing efficient. To support development and evaluation, we introduce RichDocBench for end-to-end generation, and RichDocFuzz to evaluate formatting-error recognition for reviewer agents. Through extensive automated evaluation and blind human-preference studies, we show that FormAct consistently outperforms strong baselines, including Codex-CLI, with particularly strong improvements in generating error-free, professional rich-format documents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy 等ICCV 2021 · 被引用 162 次
- Language-Guided Global Image Editing via Cross-Modal Cyclic MechanismWentao Jiang, Ning Xu, Jiayun Wang, Chen Gao 等ICCV 2021 · 被引用 28 次
相关 Paper
- Free your mouse! Command Large Language Models to Generate Code to Format Word DocumentsShihao Rao, Liang Li, Jiapeng Liu, Weixin Guan 等EMNLP 2024 · 被引用 1 次
- DocAgent: An Agentic Framework for Multi-Modal Long-Context Document UnderstandingLi Sun, Liu He, Shuyue Jia, Yangfan He 等EMNLP 2025 · 被引用 1 次
- FinSight: Towards Real-World Financial Deep ResearchJiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian 等ACL 2026 · 被引用 5 次
- MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon ReasoningYaorui Shi, Shugui Liu, Yu Yang, Wenyu Mao 等ICML 2026 · 被引用 13 次
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao 等ICML 2026
