Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
摘要
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/ amirbalef/is_one_layer_enough .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- Large Scale Transfer Learning for Tabular Data via Language ModelingJosh Gardner, Juan C. Perdomo, Ludwig SchmidtNeurIPS 2024 · 被引用 103 次
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 被引用 92 次
相关 Paper
- When and How Unlabeled Data Provably Improve In-Context LearningYingcong Li, Xiangyu Chang, Muti Kara, Xiaofeng Liu 等NeurIPS 2025 · 被引用 5 次
- SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early ExitSi-Yang Liu, Han-Jia YeICML 2026
- Universal Redundancies in Time Series Foundation ModelsAnthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William GilpinICML 2026 · 被引用 2 次
- End-to-End Compression for Tabular Foundation ModelsGuri Zabërgja, Rafiq Kamel, Arlind Kadra, Christian Frey 等ICML 2026 · 被引用 4 次
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
