Predicting Through Generation: Why Generation Is Better for Prediction
Md. Kowsher, Nusrat Jahan Prottasha, Prakash Bhat, Chun-Nam Yu, Mojtaba Soltanalian, Ivan Garibay, Ozlem O. Garibay, Chen Chen, Niloofar Yousefi
Abstract
This paper argues that generating output tokens is more effective than using pooled representations for prediction tasks because token-level generation retains more mutual information. Since LLMs are trained on massive text corpora using next-token prediction, generation aligns naturally with their learned behavior. Using the Data Processing Inequality (DPI), we provide both theoretical and empirical evidence supporting this claim. However, autoregressive models face two key challenges when used for prediction: (1) exposure bias, where the model sees ground-truth tokens during training but relies on its own predictions during inference, leading to errors, and (2) format mismatch, where discrete tokens do not always align with the task's required output structure. To address these challenges, we introduce PredGen (Predicting Through Generating), an end-to-end framework that (i) uses scheduled sampling to reduce exposure bias, and (ii) introduces a task adapter to convert the generated tokens into structured outputs. Additionally, we introduce Writer-Director Alignment Loss (WDAL), which ensures consistency between token generation and final task predictions, improving both text coherence and numerical accuracy. We evaluate PredGen on multiple classification and regression benchmarks. Our results show that PredGen consistently outperforms standard baselines, demonstrating its effectiveness in structured prediction tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c69d0b4-6192-4a4e-bfb8-10f2987bec29Cited by top-tier papers2
- LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task LearningMd Kowsher, Haris Mansoor, Nusrat Prottasha, Ozlem Garibay et al.ICML 2026
- FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language ModelsMd Kowsher, Nusrat Jahan Prottasha, Shiyun Xu, Shetu Mohanto et al.ICLR 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.ICCV 2025 · 8 citations
- Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?Tianxing He, Jingzhao Zhang, Zhiming Zhou, James R. GlassEMNLP 2021 · 12 citations
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive GenerationXinyao Liao, QIYUAN HE, Kai Xu, Xiaoye Qu et al.CVPR 2026 · 6 citations
- Rethinking Visual Autoregressive Sampling with Information-Grounding GuidanceKy Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh, Daochang Liu et al.ICML 2026
- Normalized Rewards for Preference OptimizationShawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald et al.ICML 2026 · 571 citations
