Contrastive Reinforcement Learning of Symbolic Reasoning Domains
Gabriel Poesia, Wenxin Dong, Noah D. Goodman
摘要
Abstract symbolic reasoning, as required in domains such as mathematics and logic, is a key component of human intelligence. Solvers for these domains have important applications, especially to computer-assisted education. But learning to solve symbolic problems is challenging for machine learning algorithms. Existing models either learn from human solutions or use hand-engineered features, making them expensive to apply in new domains. In this paper, we instead consider symbolic domains as simple environments where states and actions are given as unstructured text, and binary rewards indicate whether a problem is solved. This flexible setup makes it easy to specify new domains, but search and planning become challenging. We introduce four environments inspired by the Mathematics Common Core Curriculum, and observe that existing Reinforcement Learning baselines perform poorly. We then present a novel learning algorithm, Contrastive Policy Learning (ConPoLe) that explicitly optimizes the InfoNCE loss, which lower bounds the mutual information between the current state and next states that continue on a path to the solution. ConPoLe successfully solves all four domains. Moreover, problem representations learned by ConPoLe enable accurate prediction of the categories of problems in a real mathematics curriculum. Our results suggest new directions for reinforcement learning in symbolic domains, as well as applications to mathematics education.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi 等AAAI 2024 · 被引用 53 次
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic LensXixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li 等NeurIPS 2025 · 被引用 44 次
- Assistive Teaching of Motor Control Tasks to HumansMegha Srivastava, Erdem Biyik, Suvir Mirchandani, Noah D. Goodman 等NeurIPS 2022 · 被引用 12 次
- RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment TreesTengxiao Liu, Qipeng Guo, Xiangkun Hu, Yue Zhang 等EMNLP 2022 · 被引用 8 次
- Contrastive Representations for Temporal ReasoningAlicja Ziarko, Michal Bortkiewicz, Michal Zawalski, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper5
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Conditional Negative Sampling for Contrastive Learning of Visual RepresentationsMike Wu, Milan Mossé, Chengxu Zhuang, Daniel Yamins 等ICLR 2021 · 被引用 89 次
- An Interaction Design for Machine Teaching to Develop AI TutorsDaniel Weitekamp III, Erik Harpstead, Kenneth R. KoedingerCHI 2020 · 被引用 69 次
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningHaotian Fu, Hongyao Tang, Jianye Hao, Chen Chen 等AAAI 2021 · 被引用 61 次
相关 Paper
- GRACE: Generative Representation Learning via Contrastive Policy OptimizationJiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong 等ICLR 2026 · 被引用 7 次
- Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsYifan Zhou, Sachin Grover, Mohamed El Mistiri, Kamalesh Kalirathinam 等NeurIPS 2025 · 被引用 3 次
- From Imitation to Discrimination: Toward a Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning TasksChangpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang 等AAAI 2026
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu 等ICML 2026 · 被引用 102 次
- ContraBAR: Contrastive Bayes-Adaptive Deep RLEra Choshen, Aviv TamarICML 2023 · 被引用 10 次
