Don't Take This Out of Context!: On the Need for Contextual Models and Evaluations for Stylistic Rewriting
Akhila Yerukola, Xuhui Zhou, Elizabeth Clark, Maarten Sap
Abstract
Most existing stylistic text rewriting methods and evaluation metrics operate on a sentence level, but ignoring the broader context of the text can lead to preferring generic, ambiguous, and incoherent rewrites. In this paper, we investigate integrating the preceding textual context into both the rewriting and evaluation stages of stylistic text rewriting, and introduce a new composite contextual evaluation metric CtxSimFit that combines similarity to the original sentence with contextual cohesiveness. We comparatively evaluate non-contextual and contextual rewrites in formality, toxicity, and sentiment transfer tasks. Our experiments show that humans significantly prefer contextual rewrites as more fitting and natural over non-contextual ones, yet existing sentence-level automatic metrics (e.g., ROUGE, SBERT) correlate poorly with human preferences (𝜌=0–0.3). In contrast, human preferences are much better reflected by both our novel CtxSimFit (𝜌=0.7–0.9) as well as proposed context-infused versions of common metrics (𝜌=0.4–0.7). Overall, our findings highlight the importance of integrating context into the generation and especially the evaluation stages of stylistic text rewriting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e867c65a-40dd-4a18-85a2-21f6b5ce403eCited by top-tier papers2
- Words Like Knives: Backstory-Personalized Modeling and Detection of Violent CommunicationJocelyn J. Shen, Akhila Yerukola, Xuhui Zhou, Cynthia Breazeal et al.EMNLP 2025 · 4 citations
- Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeICDE 2026 · 1 citation
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 121 citations
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy et al.ACL 2022 · 96 citations
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
Related papers
- Text Detoxification using Large Pre-trained Neural ModelsDavid Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva et al.EMNLP 2021 · 16 citations
- Contextualized Rewriting for Text SummarizationGuangsheng Bao, Yue ZhangAAAI 2021 · 17 citations
- Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation FrameworkAnna Wegmann, Dong NguyenEMNLP 2021 · 7 citations
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 9 citations
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
