OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu, Runnan Fang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang
Abstract
Machine writing with large language models often relies on retrieval-augmented generation. However, these approaches remain confined within the boundaries of the model's predefined scope, limiting the generation of content with rich information. Specifically, vanillaretrieved information tends to lack depth, novelty, and suffers from redundancy, which negatively impacts the quality of generated articles, leading to shallow, unoriginal, and repetitive outputs. To address these issues, we propose OmniThink, a slow-thinking machine writing framework that emulates the human-like process of iterative expansion and reflection. The core idea behind OmniThink is to simulate the cognitive behavior of learners as they slowly deepen their knowledge of the topics. Experimental results demonstrate that OmniThink improves the knowledge density of generated articles without compromising metrics such as coherence and depth. Human evaluations and expert feedback further highlight the potential of OmniThink to address real-world challenges in the generation of long-form articles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian et al.NeurIPS 2025 · 354 citations
- WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article GenerationZhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei et al.ICCV 2025 · 3 citations
- LogiStory: A Logic-Aware Framework for Multi-Image Story VisualizationChutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao et al.ICLR 2026 · 3 citations
- Robust LLM Unlearning via Post Judgment and Multi-round ThinkingXinrui Chen, Xu Cao, Jianhao Zhang, Pinlong Zhao et al.ICLR 2026
- R4: Nested Reasoning-Retrieval for Reward Modeling in Role-Playing AgentsRenzhi Wang, Chongqiang Wei, Zhisheng Wang, Piji LiICLR 2026
Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang et al.NeurIPS 2024 · 151 citations
- RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language ModelsPeng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu et al.EMNLP 2024 · 39 citations
Related papers
- Analysis of Plan-based Retrieval for Grounded Text GenerationAmeya Godbole, Nicholas Monath, Seungyeon Kim, Ankit Singh Rawat et al.EMNLP 2024
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng et al.ACL 2026 · 12 citations
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
- IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural ThinkingZechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li et al.ACL 2026
- Does Writing with Language Models Reduce Content Diversity?Vishakh Padmakumar, He HeICLR 2024 · 173 citations
