Precise Information Control in Long-Form Text Generation
Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Zhiyuan Zeng, Weijia Shi, Yulia Tsvetkov, Danqi Chen, Pang Wei Koh, Luke Zettlemoyer
Abstract
A central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model's ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability-improving from 69.1% to 91.0% F 1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation.
Task Name PIC Type N C Example Instruction I EntityBiosPIC Full 183 50.5 Generate a factual biography about Suthida. PopBios-PPIC Full 111 20.1 Give me a biography on Erwin Schrödinger, the scientist who discovered Quantum Mechanics, Schrödinger's Cat Thought Experiment. PopBios-CFPIC Full 111 20.1 Give me a biography on Oscar Wilde, the scientist who discovered Quantum Mechanics, Schrödinger's Cat Thought Experiment. ELI5PIC Full 146 12.5 Answer the following question(s): why it's common to have 87-octane gasoline in the US but it's almost always 95-octane in Europe? AskHistoriansPIC Full 158 19.2 In the original Star Wars: A New Hope, Obi-Wan Kenobi instructs R2-D2 to connect to the Imperial network to gain access to the whole system. Did the concept of an interconnected vast computer network exist in 1977? ExpertQA PIC Full 152 13.5 Answer the question(s): What's the difference between modern and contemporary architecture? FACTSPIC Partial 150 63.5 Explain the benefits of using mobile technology to improve healthcare management in both hi-income and low-income countries. Context p XSUMPIC Partial 200 30.6 Summarize the following text in around 20-25 words. Context p
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 686ed5aa-4151-43fb-9028-9fb90f2e788aCited by top-tier papers3
- Characterizing Deep Research: A Benchmark and Formal DefinitionAbhinav Java, Ashmit Khandelwal, Sukruta Prakash Midigeshi, Aaron Halfaker et al.ICLR 2026 · 30 citations
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form GenerationFangda Ye, Kuicai Dong, Zhifei Xie, Yuxin Hu et al.ACL 2026 · 2 citations
- IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural ThinkingZechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li et al.ACL 2026
Builds on59
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text GenerationSong Duong, Florian Le Bronnec, Alexandre Allauzen, Vincent Guigue et al.ICLR 2025
- ReFF: Reinforcing Format Faithfulness in Language Models Across Varied TasksJiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen et al.AAAI 2025 · 1 citation
- Guidance: Sentence-Level Citation Enforcement via Prefix-Tail Guidance during LLM DecodingYirui Zhan, Xu, Jun GaoICML 2026
- CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language ModelsXiaqiang Tang, Jian Li, Keyu Hu, Nan Du et al.ACL 2025 · 3 citations
- Copy-Paste to Mitigate Large Language Model HallucinationsYongchao Long, Yingying Zhang, Xianbin Wen, Xian Wu et al.ICLR 2026 · 2 citations
