Filtered-CoPhy: Unsupervised Learning of Counterfactual Physics in Pixel Space
Steeven Janny, Fabien Baradel, Natalia Neverova, Madiha Nadri, Greg Mori, Christian Wolf
Abstract
Learning causal relationships in high-dimensional data (images, videos) is a hard task, as they are often defined on low dimensional manifolds and must be extracted from complex signals dominated by appearance, lighting, textures and also spurious correlations in the data. We present a method for learning counterfactual reasoning of physical processes in pixel space, which requires the prediction of the impact of interventions on initial conditions. Going beyond the identification of structural relationships, we deal with the challenging problem of forecasting raw video over long horizons. Our method does not require the knowledge or supervision of any ground truth positions or other object or scene properties. Our model learns and acts on a suitable hybrid latent representation based on a combination of dense features, sets of 2D keypoints and an additional latent vector per keypoint. We show that this better captures the dynamics of physical processes than purely dense or sparse representations. We introduce a new challenging and carefully designed counterfactual benchmark for predictions in pixel space and outperform strong baselines in physics-inspired ML and video prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d89933e-93e4-43b5-9073-95e21066dd1bCited by top-tier papers4
- Learning Physics Constrained Dynamics Using AutoencodersTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeNeurIPS 2022 · 39 citations
- Space and time continuous physics simulation from partial observationsSteeven Janny, Madiha Nadri, Julie Digne, Christian WolfICLR 2024 · 10 citations
- Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly DetectionWenqiao Li, Yao Gu, Xintao Chen, Xiaohao Xu et al.CVPR 2025
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas et al.CVPR 2025
Builds on6
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Augmenting Physical Models with Deep Networks for Complex Dynamics ForecastingYuan Yin, Vincent Le Guen, Jérémie Donà, Emmanuel de Bézenac et al.ICLR 2021 · 165 citations
- Causal Discovery in Physical Systems from VideosYunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox et al.NeurIPS 2020 · 133 citations
- CoPhy: Counterfactual Learning of Physical DynamicsFabien Baradel, Natalia Neverova, Julien Mille, Greg Mori et al.ICLR 2020 · 105 citations
Related papers
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi et al.NeurIPS 2023 · 26 citations
- Grounding Physical Concepts of Objects and Events Through Dynamic Visual ReasoningZhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong et al.ICLR 2021 · 13 citations
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
- Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual PredictionZongzhao Li, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangAAAI 2022 · 6 citations
- Weakly supervised causal representation learningJohann Brehmer, Pim de Haan, Phillip Lippe, Taco S. CohenNeurIPS 2022 · 196 citations
