Language modeling via stochastic processes
Rose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori Hashimoto
Abstract
Modern language models can generate high-quality short texts. However, they often meander or are incoherent when generating longer texts. These issues arise from the next-token-only language modeling objective. Recent work in selfsupervised learning suggests that models can learn good latent representations via contrastive learning, which can be effective for discriminative tasks. Our work analyzes the application of contrastive representations for generative tasks, like long text generation. We propose one approach for leveraging constrastive representations, which we call Time Control (TC). TC first learns a contrastive representation of the target text domain, then generates text by decoding from these representations. Compared to domain-specific methods and fine-tuning GPT2 across a variety of text domains, TC performs competitively to methods specific for learning sentence representations on discourse coherence. On long text generation settings, TC preserves the text structure both in terms of ordering (up to +15% better) and text length consistency (up to +90% better) 1 . 1 Please find our code at https://github.com/rosewang2008/language_modeling_via_ stochastic_processes Correction note This is a revised version of the original ICLR 2022 paper. During post-publication code review, we discovered that the original version of the code did not leverage goal-directedness during decoding. While we still find that contrastive representations lead to gains in our evaluations, this error affects other claims made in the paper on goal-directed decoding. To correct this and help improve our understanding of goal-directed decoding, this updated version of the manuscript contains results on both goal-directed and non goal-directed baselines. We detail the difference between the original work and this updated work in Appendix H.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f065ab2-09e6-47cf-8c7c-47eaa0763042Cited by top-tier papers8
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified FlowXingchao Liu, Chengyue Gong, Qiang LiuICLR 2023 · 75 citations
- Exploring Temporal Concurrency for Video-Language Representation LearningHeng Zhang, Daqing Liu, Zezhong Lv, Bing Su et al.ICCV 2023 · 6 citations
- Brownian Bridge Augmented Surrogate Simulation and Injection Planning for Geological CO2 StorageHaoyue Bai, Guodong Chen, Wangyang Ying, Xinyuan Wang et al.AAAI 2026 · 5 citations
- Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationZesen Cheng, Kehan Li, Hao Li, Peng Jin et al.AAAI 2025 · 3 citations
- DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn ConversationsAng Lv, Jinpeng Li, Yuhan Chen, Gao Xing et al.ACL 2023 · 3 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 264 citations
Related papers
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
- PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationZhe Hu, Hou Pong Chan, Jiachen Liu, Xinyan Xiao et al.ACL 2022
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang et al.AAAI 2026 · 3 citations
- Genre-Controllable Story Generation via Supervised Contrastive LearningJin-Uk Cho, Min-Su Jeong, JinYeong Bak, Yun-Gyung CheongWWW 2022 · 6 citations
- ContraCLM: Contrastive Learning For Causal Language ModelNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang et al.ACL 2023 · 4 citations
