The Perception–Physics Paradox: Probing Scientific Alignment with TC-Bench
Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller, Francesco Locatello
摘要
While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making even perception-based out-of-distribution accuracy a poor proxy for scientific utility. As a result, models may look correct without reasoning correctly-a discrepancy we term the Perception-Physics Paradox. To address this gap, we introduce scientific alignment as an implicit objective for representation learning in scientific domains. We study a principled, testable aspect of scientific alignment through structural isomorphism, which requires latent representations to uniquely identify physical systems up to a linear reparameterization. This perspective induces a hierarchy of necessary conditions and yields a systematic probing protocol for physical and causal interpretability. To operationalize this framework, we release TC-BENCH, a global, reproducible benchmark dataset with an automated construction pipeline for tropical cyclone research, and show that current VFMs rely on visual shortcuts that collapse in intense regimes, indicating that scientific alignment does not arise as a natural byproduct of scaling alone. We provide all evaluation code and the reproducible data construction pipeline at https://github.com/CausalLearningAI/tc-bench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele 等NeurIPS 2023 · 被引用 127 次
- Identifiability Guarantees for Causal Disentanglement from Soft InterventionsJiaqi Zhang, Kristjan H. Greenewald, Chandler Squires, Akash Srivastava 等NeurIPS 2023 · 被引用 120 次
相关 Paper
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
- OmniArch: Building Foundation Model for Scientific ComputingTianyu Chen, Haoyi Zhou, Ying Li, Hao Wang 等ICML 2025
- Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal InterventionZhe Xu, Zhicai Wang, Junkang Wu, Jinda Lu 等AAAI 2026
- Towards Multimodal Data-Driven Scientific Discovery Powered by LLM AgentsFan Liu, Xiaozhao Zeng, Hao LiuICLR 2026
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World ModelsKeyon Vafa, Peter G. Chang, Ashesh Rambachan, Sendhil MullainathanICML 2025
