Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
Chengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao, Fuli Feng, Xiangnan He, Tat-Seng Chua
Abstract
Preference alignment has enabled large language models (LLMs) to better reflect human expectations, but current methods mostly optimize for population-level preferences, overlooking individual users. Personalization is essential, yet early approaches-such as prompt customization or fine-tuning-struggle to reason over implicit preferences, limiting real-world effectiveness. Recent"think-then-generate"methods address this by reasoning before response generation. However, they face challenges in long-form generation: their static one-shot reasoning must capture all relevant information for the full response generation, making learning difficult and limiting adaptability to evolving content. To address this issue, we propose FlyThinker, an efficient"think-while-generating"framework for personalized long-form generation. FlyThinker employs a separate reasoning model that generates latent token-level reasoning in parallel, which is fused into the generation model to dynamically guide response generation. This design enables reasoning and generation to run concurrently, ensuring inference efficiency. In addition, the reasoning model is designed to depend only on previous responses rather than its own prior outputs, which preserves training parallelism across different positions-allowing all reasoning tokens for training data to be produced in a single forward pass like standard LLM training, ensuring training efficiency. Extensive experiments on real-world benchmarks demonstrate that FlyThinker achieves better personalized generation while keeping training and inference efficiency. Our code is available at https://github.com/wcb0219-sketch/FlyThinker.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e707dade-b0b2-4cdd-bf76-439420a304b1Cited by top-tier papers5
- VINCIE: Unlocking In-context Image Editing from VideoLeigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao et al.ICLR 2026 · 18 citations
- TTOM: Test-Time Optimization and Memorization for Compositional Video GenerationLeigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang et al.ICLR 2026 · 6 citations
- Intuition-Guided Latent Reasoning for LLM-Based RecommendationChang Liu, Yimeng Bai, Xiaoyan Zhao, Yang Zhang et al.KDD 2026 · 2 citations
- AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text GenerationGuanran Luo, Wentao Qiu, Wanru Zhao, Wenhan Lv et al.ACL 2026 · 2 citations
- Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesYu Wang, Emmanuele Chersoni, Chu-Ren HuangACL 2026
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Harnessing Large Language Models for Text-Rich Sequential RecommendationZhi Zheng, Wenshuo Chao, Zhaopeng Qiu, Hengshu Zhu et al.WWW 2024 · 114 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
Related papers
- OneRec-Think: In-Text Reasoning for Generative RecommendationZhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang et al.ACL 2026 · 48 citations
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsIlgee Hong, Changlong Yu, Liang Qiu, Weixiang Yan et al.NeurIPS 2025 · 15 citations
- LightThinker: Thinking Step-by-Step CompressionJintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo et al.EMNLP 2025 · 2 citations
- NextQuill: Causal Preference Modeling for Enhancing LLM PersonalizationXiaoyan Zhao, Juntao You, Yang Zhang, Wenjie Wang et al.ICLR 2026 · 38 citations
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho et al.ICML 2026 · 7 citations
