DocEdit: Language-Guided Document Editing
Puneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt, Dinesh Manocha, Vlad I. Morariu
Abstract
Professional document editing tools require a certain level of expertise to perform complex edit operations. To make editing tools accessible to increasingly novice users, we investigate intelligent document assistant systems that can make or suggest edits based on a user's natural language request. Such a system should be able to understand the user's ambiguous requests and contextualize them to the visual cues and textual content found in a document image to edit localized unstructured text and structured layouts. To this end, we propose a new task of language-guided localized document editing, where the user provides a document and an open vocabulary editing request, and the intelligent system produces a command that can be used to automate edits in real-world document editing software. In support of this task, we curate the DocEdit dataset, a collection of approximately 28K instances of user edit requests over PDF and design templates along with their corresponding ground truth software executable commands. To our knowledge, this is the first dataset that provides a diverse mix of edit operations with direct and indirect references to the embedded text and visual objects such as paragraphs, lists, tables, etc. We also propose DocEditor, a Transformer-based localization-aware multimodal (textual, spatial, and visual) model that performs the new task. The model attends to both document objects and related text contents which may be referred to in a user edit request, generating a multimodal embedding that is used to predict an edit command and associated bounding box localizing it. Our proposed model empirically outperforms other baseline deep learning approaches by 15-18%, providing a strong starting point for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e68fd105-b66d-4251-a638-93e30fe97eb6Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACM MM 2022 · 184 citations
- Form2Seq : A Framework for Higher-Order Form Structure ExtractionMilan Aggarwal, Hiresh Gupta, Mausoom Sarkar, Balaji KrishnamurthyEMNLP 2020 · 18 citations
- SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color EditingJing Shi, Ning Xu, Haitian Zheng, Alex Smith et al.CVPR 2022 · 15 citations
Related papers
- DocEdit-v2: Document Structure Editing Via Multimodal LLM GroundingManan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain et al.EMNLP 2024 · 1 citation
- OBJECT 3DIT: Language-guided 3D-aware Image EditingOscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna et al.NeurIPS 2023 · 79 citations
- M3L: Language-based Video Editing via Multi-Modal Multi-Level TransformersTsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein et al.CVPR 2022 · 13 citations
- Multimodal Markup Document Models for Graphic Design CompletionKotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani et al.ACM MM 2025 · 1 citation
- InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with InstructionsRyota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito et al.AAAI 2024 · 39 citations
