Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLP
Zhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schölkopf
Abstract
The principle of independent causal mechanisms (ICM) states that generative processes of real world data consist of independent modules which do not influence or inform each other. While this idea has led to fruitful developments in the field of causal inference, it is not widely-known in the NLP community. In this work, we argue that the causal direction of the data collection process bears nontrivial implications that can explain a number of published NLP findings, such as differences in semi-supervised learning (SSL) and domain adaptation (DA) performance across different settings. We categorize common NLP tasks according to their causal direction and empirically assay the validity of the ICM principle for text data using minimum description length. We conduct an extensive meta-analysis of over 100 published SSL and 30 DA studies, and find that the results are consistent with our expectations based on causal insights. This work presents the first attempt to analyze the ICM principle in NLP, and provides constructive suggestions for future modeling choices. 1 * Equal contribution. 1 The codes are at https://github.com/zhijing-jin/icm4nlp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- CauSSL: Causality-inspired Semi-supervised Learning for Medical Image SegmentationJuzheng Miao, Cheng Chen, Furui Liu, Hao Wei et al.ICCV 2023 · 88 citations
- A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language ModelsAlessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf et al.ACL 2023 · 15 citations
- Causal vs. Anticausal merging of predictorsSergio Hernan Garrido Mejia, Patrick Blöbaum, Bernhard Schölkopf, Dominik JanzingNeurIPS 2024 · 1 citation
- Sequential Learning of Neural Networks for Prequential MDLJörg Bornschein, Yazhe Li, Marcus HutterICLR 2023 · 1 citation
- Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng et al.ACL 2025
Builds on6
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- Discovering Fully Oriented Causal NetworksOsman Mian, Alexander Marx, Jilles VreekenAAAI 2021 · 37 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal EstimatesKatherine A. Keith, David D. Jensen, Brendan O'ConnorACL 2020 · 16 citations
- On The Evaluation of Machine Translation SystemsTrained With Back-TranslationSergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael AuliACL 2020 · 15 citations
Related papers
- Can Large Language Models Learn Independent Causal Mechanisms?Gaël Gendron, Bao Trung Nguyen, Alex Yuxuan Peng, Michael J. Witbrock et al.EMNLP 2024 · 2 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- Diverse Distributions of Self-Supervised Tasks for Meta-Learning in NLPTrapit Bansal, Karthick Prasad Gunasekaran, Tong Wang, Tsendsuren Munkhdalai et al.EMNLP 2021 · 27 citations
- Cross-Lingual Transfer with Class-Weighted Language-Invariant RepresentationsRuicheng Xian, Heng Ji, Han ZhaoICLR 2022 · 5 citations
- iTAG: Inverse Design for Natural Text Generation with Accurate Causal Graph AnnotationsWenshuo Wang, Boyu Cao, Nan Zhuang, Wei LiACL 2026
