Deep Natural Language Feature Learning for Interpretable Prediction
Felipe Urrutia, Cristian Buc Calderon, Valentin Barrière
摘要
We propose a general method to break down a main complex task into a set of intermediary easier sub-tasks, which are formulated in natural language as binary questions related to the final target task. Our method allows for representing each example by a vector consisting of the answers to these questions. We call this representation Natural Language Learned Features (NLLF). NLLF is generated by a small transformer language model (e.g., BERT) that has been trained in a Natural Language Inference (NLI) fashion, using weak labels automatically obtained from a Large Language Model (LLM). We show that the LLM normally struggles for the main task using in-context learning, but can handle these easiest subtasks and produce useful weak labels to train a BERT. The NLI-like training of the BERT allows for tackling zeroshot inference with any binary question, and not necessarily the ones seen during the training. We show that this NLLF vector not only helps to reach better performances by enhancing any classifier, but that it can be used as input of an easy-to-interpret machine learning model like a decision tree. This decision tree is interpretable but also reaches high performances, surpassing those of a pre-trained transformer in some cases. We have successfully applied this method to two completely different tasks: detecting incoherence in students' answers to open-ended mathematics exam questions, and screening abstracts for a systematic literature review of scientific papers on climate change and agroecology. 1 Interpretable ML model Q: Does the article address the relationship between agroecological practices and climate change? A: Yes The text has a word with the prefix convent Q: Does the article analyze how agroecology affects nitrogen dynamics? A: Yes Include ✔ Q: Does the article assess agroecological practices' impact on climate change? A: Yes Contribution of crop residue, soil, and fertilizer nitrogen to nitrous oxide emissions varies with long-term crop rotation and tillage Agriculture is an important contributor to N2O emissions -a potent greenhouse gas -with high peaks occurring when soil mineral nitrogen (N) is high (e.g., after mineralization of organic N and N fertilizer application). Nitrogen dynamics in soil and consequently N2O emissions are affected by crop and soil management practices (e.g., crop rotation and tillage), an effect mostly assessed in the literature through comparisons of total N2O emission. Hence, information is scarce on the effect of these management practices on specific N sources affecting N2O emissions (i.e., N fertilizer, soil, above and belowground crop residues) -a knowledge gap explored in this study with the use of N-15 tracers. The isotope approach enabled … (more)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu 等ICLR 2023 · 被引用 94 次
- Rule Discovery for Natural Language Inference Data Generation Using Out-of-Distribution DetectionJuyoung Han, Hyunsun Hwang, Changki LeeEMNLP 2025
- Making Text Embedders Few-Shot LearnersChaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen 等ICLR 2025
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello 等NeurIPS 2024 · 被引用 26 次
