Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive Feedback
Qian Dong, Yiding Liu, Qingyao Ai, Zhijing Wu, Haitao Li, Yiqun Liu, Shuaiqiang Wang, Dawei Yin, Shaoping Ma
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities across various research domains, including the field of Information Retrieval (IR). However, the responses generated by off-the-shelf LLMs tend to be generic, i.e., cannot capture the distinctiveness of each document with similar content. This limits the performance of LLMs in IR because finding and distinguishing relevant documents from substantial similar documents is a typical problem in many IR tasks. To address this issue, we propose an unsupervised alignment method, namely Reinforcement Learning from Contrastive Feedback (RLCF), empowering LLMs to generate both high-quality and context-specific responses. Our approach constructs unsupervised contrastive feedback signals based on similar document groups, and adopts a reward function, named group-wise reciprocal rank, to optimize LLMs within a standard Proximal Policy Optimization. We conduct extensive experiments to evaluate the effectiveness of RLCF on LLMs built with different languages and parameter sizes on multiple downstream IR applications. RLCF significantly outperforms existing alignment methods, and RLCF-optimized LLMs demonstrate considerable improvement in generating responses with distinctiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecbf435c-224e-4c87-9872-1e1d8b63d2e2Cited by top-tier papers6
- Enhancing Uncertainty Modeling with Semantic Graph for Hallucination DetectionKedi Chen, Qin Chen, Jie Zhou, Xinqi Tao et al.AAAI 2025 · 13 citations
- Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative DecodingYunjia Xi, Hangyu Wang, Bo Chen, Jianghao Lin et al.SIGIR 2025 · 5 citations
- DMAP: Human-Aligned Structural Document Map for Multimodal Document UnderstandingShunliang Fu, Yanxin Zhang, Yixin Xiang, Xiaoyu Du et al.WWW 2026
- CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-JudgesHaitao Li, Junjie Chen, Qingyao Ai, Zhumin Chu et al.ACL 2025
- Modular Representation Compression: Adapting LLM Representations for Efficient and Effective RecommendationYunjia Xi, Menghui Zhu, Jianghao Lin, Bo Chen et al.SIGIR 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
Related papers
- Aligning Large Language Models for Controllable RecommendationsWensheng Lu, Jianxun Lian, Wei Zhang, Guanghua Li et al.ACL 2024 · 7 citations
- On a Connection Between Imitation Learning and RLHFTeng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen et al.ICLR 2025
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu et al.ACL 2025
- Reinforcement Learning for Large Language Models via Group Preference Reward ShapingHuaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao et al.EMNLP 2025
- Enhancing Reinforcement Learning with Label-Sensitive Reward for Natural Language UnderstandingKuo Liao, Shuang Li, Meng Zhao, Liqun Liu et al.ACL 2024 · 2 citations
