Discovery of the Hidden World with Large Language Models
Chenxi Liu, Yongqiang Chen, Tongliang Liu, Mingming Gong, James Cheng, Bo Han, Kun Zhang
摘要
Science originates with discovering new causal knowledge from a combination of known facts and observations . Traditional causal discovery approaches mainly rely on high-quality measured variables, usually given by human experts, to find causal relations. However, the causal variables are usually unavailable in a wide range of real-world applications. The rise of large language models (LLMs) that are trained to learn rich knowledge from the massive observations of the world, provides a new opportunity to assist with discovering high-level hidden variables from the raw observational data. Therefore, we introduce COAT: C ausal representati O n A ssistan T . COAT incorporates LLMs as an factor proposer that extracts the potential causal factors from unstructured data . Moreover, LLMs can also be instructed to provide additional information used to collect data values (e.g., annotation criteria) and to further parse the raw unstructured data into structured data. The annotated data will be fed to a causal learning module (e.g., the FCI algorithm) that provides both rigorous explanations of the data, as well as useful feedback to further improve the extraction of causal factors by LLMs. We verify the effectiveness of COAT in uncovering the underlying causal system with two case studies of review rating analysis and neuropathic diagnosis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh 等NeurIPS 2025 · 被引用 23 次
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?Junchi Yu, Yujie Liu, Jindong Gu, Philip H. S. Torr 等NeurIPS 2025 · 被引用 8 次
- Revealing Multimodal Causality with Large Language ModelsJin Li, Shoujin Wang, Qi Zhang, Feng Liu 等NeurIPS 2025 · 被引用 5 次
- Who You Are Matters: Bridging Interests and Social Roles via LLM-Enhanced Logic RecommendationQing Yu, Xiaobei Wang, Shuchang Liu, Yandong Bai 等NeurIPS 2025 · 被引用 4 次
- Hierarchical Graph Tokenization for Molecule-Language AlignmentYongqiang Chen, Quanming Yao, Juzheng Zhang, James Cheng 等ICML 2025 · 被引用 2 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
相关 Paper
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan 等NeurIPS 2024 · 被引用 24 次
- Traceable Latent Variable Discovery Based on Multi-Agent CollaborationHuaming Du, Tao Hu, Yijie Huang, Yu Zhao 等WWW 2026
- Multistage Feedback-Driven Causal Discovery from Textual Data with Large Language ModelsJuntao Yang, Dayuan Cao, Kui Yu, Xiang Wang 等WWW 2026
- Causal Discovery through Synergizing Large Language Model and Data-Driven ReasoningHuaming Du, Yujia Zheng, Baoyu Jing, Yu Zhao 等KDD 2025 · 被引用 1 次
- Causal Modelling Agents: Causal Graph Discovery through Synergising Metadata- and Data-driven ReasoningAhmed Abdulaal, Adamos Hadjivasiliou, Nina Montaña Brown, Tiantian He 等ICLR 2024 · 被引用 43 次
