Reducing Sequence Length by Predicting Edit Spans with Large Language Models
Masahiro Kaneko, Naoaki Okazaki
Abstract
Large Language Models (LLMs) have demonstrated remarkable performance in various tasks and gained significant attention. LLMs are also used for local sequence transduction tasks, including grammatical error correction (GEC) and formality style transfer, where most tokens in a source text are kept unchanged. However, the models that generate all target tokens in such tasks have a tendency to simply copy the input text as is, without making needed changes, because the difference between input and output texts is minimal in the training data. This is also inefficient because the computational cost grows quadratically with the target sequence length with Transformer. This paper proposes predicting edit spans for the source text for local sequence transduction tasks. Representing an edit span with a position of the source text and corrected tokens, we can reduce the length of the target sequence and the computational cost for inference. We apply instruction tuning for LLMs on the supervision data of edit spans. Experiments show that the proposed method achieves comparable performance to the baseline in four tasks, paraphrasing, formality style transfer, GEC, and text simplification, despite reducing the length of the target text by as small as 21%. Furthermore, we report that the task-specific fine-tuning with the proposed method achieved state-of-the-art performance in the four tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65334d81-d611-481c-9999-17f3a49ce0aaCited by top-tier papers4
- Sprout: Green Generative AI with Carbon-Efficient LLM InferenceBaolin Li, Yankai Jiang, Vijay Gadepally, Devesh TiwariEMNLP 2024 · 11 citations
- Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection CapabilitiesMasahiro Kaneko, Timothy BaldwinEMNLP 2025 · 2 citations
- Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case StudyBashar Alhafni, Nizar HabashACL 2025
- CoAM: Corpus of All-Type Multiword ExpressionsYusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman et al.ACL 2025
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
Related papers
- Seq2Edits: Sequence Transduction Using Span-level Edit OperationsFelix Stahlberg, Shankar KumarEMNLP 2020 · 1 citation
- Monolingual Transfer Learning via Bilingual Translators for Style-Sensitive Paraphrase GenerationTomoyuki Kajiwara, Biwa Miura, Yuki AraseAAAI 2020 · 8 citations
- Whose Instructions Count? Resolving Preference Bias in Instruction Fine-TuningJiayu Zhang, Changbang Li, Yinan Peng, Weihao Luo et al.NeurIPS 2025 · 2 citations
- The Missing Alignment Link of In-context Learning on SequencesHarshvardhan Agarwal, Sunita SarawagiICML 2025
- Improving Iterative Text Revision by Learning Where to Edit from Other Revision TasksZae Myung Kim, Wanyu Du, Vipul Raheja, Dhruv Kumar et al.EMNLP 2022 · 8 citations
