InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron C. Wallace, Junyi Jessy Li
Abstract
Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes INFOLOSSQA, a framework to characterize and recover simplificationinduced information loss in form of questionand-answer (QA) pairs. Building on the theory of Questions Under Discussion, the QA pairs are designed to help readers deepen their knowledge of a text. First, we collect a dataset of 1,000 linguist-curated QA pairs derived from 104 LLM simplifications of English medical study abstracts. Our analyses of this data reveal that information loss occurs frequently, and that the QA pairs give a high-level overview of what information was lost. Second, we devise two methods for this task: end-to-end prompting of open-source and commercial language models, and a natural language inference pipeline. With a novel evaluation framework considering the correctness of QA pairs and their linguistic suitability, our expert evaluation reveals that models struggle to reliably identify information loss and applying similar standards as humans at what constitutes information loss. 1 * Work done while visiting UT Austin. 1 Code, dataset and an interactive data viewer is available at https://jantrienes.github.io/ts-info-loss/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- FactPICO: Factuality Evaluation for Plain Language Summarization of Medical EvidenceSebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Göke et al.ACL 2024 · 12 citations
- Which questions should I answer? Salience Prediction of Inquisitive QuestionsYating Wu, Ritika Mangla, Alex Dimakis, Greg Durrett et al.EMNLP 2024 · 1 citation
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- Discourse Level Factors for Sentence Deletion in Text SimplificationYang Zhong, Chao Jiang, Wei Xu, Junyi Jessy LiAAAI 2020 · 57 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- Elaborative Simplification as Implicit Questions Under DiscussionYating Wu, William Sheffield, Kyle Mahowald, Junyi Jessy LiEMNLP 2023 · 4 citations
- Evaluating Factuality in Text SimplificationAshwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy LiACL 2022
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 2 citations
- Explainable Prediction of Text Complexity: The Missing Preliminaries for Text SimplificationCristina Garbacea, Mengtian Guo, Samuel Carton, Qiaozhu MeiACL 2021
- PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&AAnna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet KaurCHI 2026 · 3 citations
