Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge Grounding
Mingkuan Zhao, Xiayu Sun, Wentao Hu, Suquan Chen, Jiaxuan Li, Xiaoyan Zhu, Xin Lai, Jiayin Wang
Abstract
Causal language models factorize sequence probabilities using only preceding context, leaving future information unexploited during training despite its availability in the training data. This paper introduces Regret Pre-training, a selfsupervised framework grounded in the Learning Using Privileged Information (LUPI) paradigm. The framework employs a dual-view architecture in which a single model generates both a causal Student distribution and a future-conditioned Teacher distribution. The training objective augments standard language modeling with a regret loss that minimizes the KL divergence from teacher to student, transferring future-aware signals to the causal representations. We investigate two teacher configurations on the OLMoE-1B-7B architecture: LOCALREGRET, which extends attention by one future token, and GLOBALRE-GRET, which conditions on bidirectional context with the target position masked. Experiments on nine downstream tasks following 4 billion tokens of training demonstrate that both configurations consistently outperform the baseline. On average, GLOBALREGRET and LOCALREGRET achieve 33.9% and 32.2% accuracy respectively, surpassing the baseline's 30.2%. Most notably, GLOBAL-REGRET improves BoolQ performance by 18.1 percentage points (61.0% vs 42.9%). The framework introduces no additional parameters and requires only one extra inference-mode forward pass per training step. The source code for this paper is publicly available at https://github. com/RegretPretraining/Code2026 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10d980a4-7116-49b0-a337-3190a03a0f80Builds on5
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen et al.ACL 2025 · 7 citations
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsWentao Hu, Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu et al.AAAI 2026 · 4 citations
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-offMingkuan Zhao, Wentao Hu, Jiayin Wang, Xin Lai et al.AAAI 2026 · 2 citations
- Awakening Dormant Experts: Counterfactual Routing to Mitigate MoE HallucinationsWentao Hu, Yanbo Zhai, Xiaohui Hu, Mingkuan Zhao et al.ACL 2026
Related papers
- FutureTOD: Teaching Future Knowledge to Pre-trained Language Model for Task-Oriented DialogueWeihao Zeng, Keqing He, Yejie Wang, Chen Zeng et al.ACL 2023 · 3 citations
- Next-ToBE: Probabilistic Next Token-Bag Exploitation for Activating Anticipatory Capacity in LLMsYihe Liu, Huibin Wang, Xianming Hu, Pinyi Zhang et al.ICLR 2026
- Fostering Video Reasoning via Next-Event PredictionHaonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du et al.ICLR 2026 · 14 citations
- Retaining by Doing: The Role of On-Policy Data in Mitigating ForgettingHoward Chen, Noam Razin, Karthik Narasimhan, Danqi ChenICML 2026
- Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual TokenAiliang Lin, Zhuoyun Li, Yusong Wang, Kotaro Funakoshi et al.ACL 2026
