How Transformers Learn Structured Data: Insights From Hierarchical Filtering
Jerome Garnier-Brun, Marc Mézard, Emanuele Moscato, Luca Saglietti
Abstract
Understanding the learning process and the embedded computation in transformers is becoming a central goal for the development of interpretable AI. In the present study, we introduce a hierarchical filtering procedure for data models of sequences on trees, allowing us to handtune the range of positional correlations in the data. Leveraging this controlled setting, we provide evidence that vanilla encoder-only transformers can approximate the exact inference algorithm when trained on root classification and masked language modeling tasks, and study how this computation is discovered and implemented. We find that correlations at larger distances, corresponding to increasing layers of the hierarchy, are sequentially included by the network during training. By comparing attention maps from models trained with varying degrees of filtering and by probing the different encoder levels, we find clear evidence of a reconstruction of correlations on successive length scales corresponding to the various levels of the hierarchy, which we relate to a plausible implementation of the exact inference algorithm within the same architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a4b5bd4-dde1-4826-a3dd-76023a93ea3dCited by top-tier papers9
- A distributional simplicity bias in the learning dynamics of transformersRiccardo Rende, Federica Gerace, Alessandro Laio, Sebastian GoldtNeurIPS 2024 · 30 citations
- Biased Generalization in Diffusion ModelsLuca Saglietti, Luca Biggio, Jerome Garnier-Brun, Davide Beltrame et al.ICML 2026 · 2 citations
- On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case StudyRiccardo Alberghi, Elizaveta Demyanenko, Luca Biggio, Luca SagliettiNeurIPS 2025 · 2 citations
- A theory of learning data statistics in diffusion models, from easy to hardLorenzo Bardone, Claudia Merger, Sebastian GoldtICML 2026
- Probing the Latent Hierarchical Structure of Data via Diffusion ModelsAntonio Sclocchi, Alessandro Favero, Noam Itzhak Levi, Matthieu WyartICLR 2025
Builds on9
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 183 citations
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 181 citations
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 58 citations
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 33 citations
- A distributional simplicity bias in the learning dynamics of transformersRiccardo Rende, Federica Gerace, Alessandro Laio, Sebastian GoldtNeurIPS 2024 · 30 citations
Related papers
- Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional EncodingDaichi Hayakawa, Issei SatoACL 2025
- Characterizing intrinsic compositionality in transformers with Tree ProjectionsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningICLR 2023 · 13 citations
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Algorithmic Capabilities of Random TransformersZiqian Zhong, Jacob AndreasNeurIPS 2024 · 23 citations
- Internal Planning in Language Models: Characterizing Horizon and Branch AwarenessMuhammed Ustaomeroglu, Baris Askin, Gauri Joshi, Carlee Joe-Wong et al.ICLR 2026
