Retrieval Augmentation for Commonsense Reasoning: A Unified Approach
Wenhao Yu, Chenguang Zhu, Zhihan Zhang, Shuohang Wang, Zhuosheng Zhang, Yuwei Fang, Meng Jiang
Abstract
A common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled. However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever. In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks. We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever. We conducted experiments on four different commonsense reasoning tasks. Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards. Our code is available at https://github.com/wyu97/RACo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge GraphJiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang et al.ICLR 2024 · 247 citations
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QAChanghao Wang, Yanfang Liu, Xinxin Fan, Ao Tian et al.ICML 2026
- UniRAG: Unified Query Understanding Method for Retrieval Augmented GenerationRui Li, Liyang He, Qi Liu, Zheng Zhang et al.ACL 2025
- ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question AnsweringFrancesco Molfese, Simone Conia, Riccardo Orlando, Roberto NavigliEMNLP 2024
- Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense ReasoningJunho Kim, Soyeon Bak, Mingyu Lee, Minju Hong et al.EMNLP 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
Related papers
- KGR4: Retrieval, Retrospect, Refine and Rethink for Commonsense GenerationXin Liu, Dayiheng Liu, Baosong Yang, Haibo Zhang et al.AAAI 2022 · 9 citations
- Improving Commonsense in Vision-Language Models via Knowledge Graph RiddlesShuquan Ye, Yujia Xie, Dongdong Chen, Yichong Xu et al.CVPR 2023
- SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeAndong Wang, Bo Wu, Sunli Chen, Zhenfang Chen et al.CVPR 2024 · 7 citations
- Metric-guided Distillation: Distilling Knowledge from the Metric to Ranker and Retriever for Generative Commonsense ReasoningXingwei He, Yeyun Gong, A-Long Jin, Weizhen Qi et al.EMNLP 2022 · 10 citations
- RuAG: Learned-rule-augmented Generation for Large Language ModelsYudi Zhang, Pei Xiao, Lu Wang, Chaoyun Zhang et al.ICLR 2025
