Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
Jihan Yao, Wenxuan Ding, Shangbin Feng, Lucy Lu Wang, Yulia Tsvetkov
Abstract
In the absence of abundant reliable annotations for challenging tasks and contexts, how can we expand the frontier of LLM capabilities with potentially wrong answers? We focus on two research questions: (1) Can LLMs generate reliable preferences among wrong options? And if so, (2) Would alignment with such wrong-over-wrong preferences be helpful? We employ methods based on selfconsistency, token probabilities, and LLM-as-a-judge to elicit wrong-over-wrong preferences, and fine-tune language models with preference optimization approaches using these synthesized preferences. Extensive experiments with seven LLMs and eight datasets demonstrate that (1) LLMs do have preliminary capability in distinguishing various shades of wrong, achieving up to 20.9% higher performance than random guess; (2) Alignment with wrong-over-wrong preferences helps LLMs to produce less wrong and sometimes even outright correct answers, while overall improving model calibration. Code and data are publicly available at https://github.com/yaojh18/Varying-Shades-of-Wrong . * equal contribution 1 Wrongness proxies are by no means perfect; they merely serve as objective and quantitative measures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ee87d8c-5c02-46ab-860f-514cc5e4036cCited by top-tier papers4
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett et al.NeurIPS 2025 · 8 citations
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error ScenariosShiting Huang, Zhen Fang, Zehui Chen, Siyu Yuan et al.EMNLP 2025
- Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional DataHyunji (Alex) Nam, Haoran Li, Natasha JaquesICML 2026
Builds on47
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
Related papers
- CREAM: Consistency Regularized Self-Rewarding Language ModelsZhaoyang Wang, Weilei He, Zhiyuan Liang, Xuchao Zhang et al.ICLR 2025
- Preference-Strength-Aware Self-Improving Alignment with Generative Preference ModelsYuanzhao Zhai, Zhuo Zhang, Cheng Yang, Kele Xu et al.SIGIR 2025
- Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for ReasoningYongqi Tong, Dawei Li, Sizhe Wang, Yujia Wang et al.ACL 2024
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 3 citations
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou et al.ACL 2024
