Embedding-Aligned Language Models
Guy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Lior Shani, Yi Liang, Craig Boutilier
Abstract
We propose a novel approach for training large language models (LLMs) to adhere to objectives defined within a latent embedding space. Our method leverages reinforcement learning (RL), treating a pre-trained LLM as an environment. Our embedding-aligned guided language (EAGLE) agent is trained to iteratively steer the LLM's generation towards optimal regions of the latent embedding space, w.r.t. some predefined criterion. We demonstrate the effectiveness of the EAGLE agent using the MovieLens 25M and Amazon Review datasets to surface content gaps that satisfy latent user demand. We also demonstrate the benefit of using an optimal design of a state-dependent action set to improve EAGLE's efficiency. Our work paves the way for controlled and grounded text generation using LLMs, ensuring consistency with domain-specific knowledge and data representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab6fde2d-a2ad-44a0-a5d0-27bf4caf86c7Cited by top-tier papers3
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin et al.NeurIPS 2025 · 7 citations
- Preference Adaptive and Sequential Text-to-Image GenerationOfir Nabati, Guy Tennenholtz, Chih-Wei Hsu, Moonkyung Ryu et al.ICML 2025
- Comparing Few to Rank Many: Active Human Preference Learning Using Randomized Frank-Wolfe MethodKiran Koshy Thekumparampil, Gaurush Hiranandani, Kousha Kalantari, Shoham Sabach et al.ICML 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
Related papers
- Aligning Large Language Models for Controllable RecommendationsWensheng Lu, Jianxun Lian, Wei Zhang, Guanghua Li et al.ACL 2024 · 7 citations
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- Re-SpS: A Reinforcement Learning Approach to Speculative SamplingChenan Wang, Daniel H. Shi, Haipeng ChenAAAI 2026
- Improving Knowledge Extraction from LLMs for Task Learning through Agent AnalysisJames R. Kirk, Robert E. Wray, Peter Lindes, John E. LairdAAAI 2024 · 8 citations
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationRajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel et al.ICLR 2023 · 54 citations
