Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
Beatriz Borges, Niket Tandon, Tanja Käser, Antoine Bosselut
摘要
Natural Language Feedback (NLF) is an increasingly popular mechanism for aligning Large Language Models (LLMs) to human preferences.Despite the diversity of the information it can convey, NLF methods are often handdesigned and arbitrary, with little systematic grounding.At the same time, research in learning sciences has long established several effective feedback models.In this opinion piece, we compile ideas from pedagogy to introduce FELT, a feedback framework for LLMs that outlines various characteristics of the feedback space, and a feedback content taxonomy based on these variables, providing a general mapping of the feedback space.In addition to streamlining NLF designs, FELT also brings out new, unexplored directions for research in NLF.We make our taxonomy available to the community, providing guides and examples for mapping our categorizations to future research.How should feedback be provided to LLMs? Certain works augment the model through data augmentation (Shi et al., 2024), external corrective
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning SignalYuhan Liu, Michael J. Q. Zhang, Eunsol ChoiEMNLP 2025 · 被引用 12 次
- Position: LLMs Can be Good Tutors in English EducationJingheng Ye, Shen Wang, Deqing Zou, Yibo Yan 等EMNLP 2025 · 被引用 2 次
- Critique-Guided Distillation for Robust Reasoning via RefinementBerkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar, Michael Lee 等ICML 2026
它引用的顶会 Paper28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin 等ACL 2026 · 被引用 35 次
- DRESS : Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language FeedbackYangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji 等CVPR 2024
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesHannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Röttger 等EMNLP 2023 · 被引用 13 次
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang 等ACL 2026
- DeAL: Decoding-time Alignment for Large Language ModelsJames Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai 等ACL 2025
