Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
Fangzhi Xu, Qiushi Sun, Kanzhi Cheng, Jun Liu, Yu Qiao, Zhiyong Wu
摘要
One of the primary driving forces contributing to the superior performance of Large Language Models (LLMs) is the extensive availability of human-annotated natural language data, which is used for alignment fine-tuning. This inspired researchers to investigate self-training methods to mitigate the extensive reliance on human annotations. However, the current success of self-training has been primarily observed in natural language scenarios, rather than in the increasingly important neural-symbolic scenarios. To this end, we propose an environment-guided neural-symbolic self-training framework named ENVISIONS. It aims to overcome two main challenges: (1) the scarcity of symbolic data, and (2) the limited proficiency of LLMs in processing symbolic language. Extensive evaluations conducted on three distinct domains demonstrate the effectiveness of our approach. Additionally, we have conducted a comprehensive analysis to uncover the factors contributing to ENVISIONS's success, thereby offering valuable insights for future research in this area. Code will be available at https://github.com/xufangzhi/ENVISIONS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingMuye Huang, Han Lai, Xinyu Zhang, Wenjun Wu 等AAAI 2025 · 被引用 30 次
- NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-NeuronsHaonan Dong, Kehan Jiang, Haoran Ye, Wenhao Zhu 等ACL 2026 · 被引用 15 次
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingMuye Huang, Lingling Zhang, Jie Ma, Han Lai 等NeurIPS 2025 · 被引用 13 次
- AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive ReasoningBowen Ping, Minnan Luo, Zhuohang Dang, Chenxi Wang 等ICLR 2026 · 被引用 12 次
- φ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and ExploitationFangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao 等ACL 2025 · 被引用 10 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
相关 Paper
- Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language ModelsFangzhi Xu, Zhiyong Wu, Qiushi Sun, Siyu Ren 等ACL 2024
- Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic VerificationChuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan 等ICML 2026 · 被引用 3 次
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task PlanningSanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park 等NeurIPS 2025 · 被引用 14 次
- Logically Consistent Language Models via Neuro-Symbolic IntegrationDiego Calanzone, Stefano Teso, Antonio VergariICLR 2025 · 被引用 2 次
- Language-based Trial and Error Falls Behind in the Era of ExperienceHaoyu Wang, Guozheng Ma, Shugang Cui, Yilun Kong 等ICML 2026 · 被引用 2 次
