Language modeling via stochastic processes
Rose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori Hashimoto
摘要
Modern language models can generate high-quality short texts. However, they often meander or are incoherent when generating longer texts. These issues arise from the next-token-only language modeling objective. Recent work in selfsupervised learning suggests that models can learn good latent representations via contrastive learning, which can be effective for discriminative tasks. Our work analyzes the application of contrastive representations for generative tasks, like long text generation. We propose one approach for leveraging constrastive representations, which we call Time Control (TC). TC first learns a contrastive representation of the target text domain, then generates text by decoding from these representations. Compared to domain-specific methods and fine-tuning GPT2 across a variety of text domains, TC performs competitively to methods specific for learning sentence representations on discourse coherence. On long text generation settings, TC preserves the text structure both in terms of ordering (up to +15% better) and text length consistency (up to +90% better) 1 . 1 Please find our code at https://github.com/rosewang2008/language_modeling_via_ stochastic_processes Correction note This is a revised version of the original ICLR 2022 paper. During post-publication code review, we discovered that the original version of the code did not leverage goal-directedness during decoding. While we still find that contrastive representations lead to gains in our evaluations, this error affects other claims made in the paper on goal-directed decoding. To correct this and help improve our understanding of goal-directed decoding, this updated version of the manuscript contains results on both goal-directed and non goal-directed baselines. We detail the difference between the original work and this updated work in Appendix H.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified FlowXingchao Liu, Chengyue Gong, Qiang LiuICLR 2023 · 被引用 75 次
- Exploring Temporal Concurrency for Video-Language Representation LearningHeng Zhang, Daqing Liu, Zezhong Lv, Bing Su 等ICCV 2023 · 被引用 6 次
- Brownian Bridge Augmented Surrogate Simulation and Injection Planning for Geological CO2 StorageHaoyue Bai, Guodong Chen, Wangyang Ying, Xinyuan Wang 等AAAI 2026 · 被引用 5 次
- Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationZesen Cheng, Kehan Li, Hao Li, Peng Jin 等AAAI 2025 · 被引用 3 次
- DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn ConversationsAng Lv, Jinpeng Li, Yuhan Chen, Gao Xing 等ACL 2023 · 被引用 3 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 被引用 264 次
相关 Paper
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationZhe Hu, Hou Pong Chan, Jiachen Liu, Xinyan Xiao 等ACL 2022
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang 等AAAI 2026 · 被引用 3 次
- Genre-Controllable Story Generation via Supervised Contrastive LearningJin-Uk Cho, Min-Su Jeong, JinYeong Bak, Yun-Gyung CheongWWW 2022 · 被引用 6 次
- ContraCLM: Contrastive Learning For Causal Language ModelNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 等ACL 2023 · 被引用 4 次
