Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional Encoding
Daichi Hayakawa, Issei Sato
Abstract
In this study, we provide constructive proof that Transformers can recognize and generate hierarchical language efficiently with respect to model size, even without the need for a specific positional encoding. Specifically, we show that causal masking and a starting token enable Transformers to compute positional information and depth within hierarchical structures. We demonstrate that Transformers without positional encoding can generate hierarchical languages. Furthermore, we suggest that explicit positional encoding might have a detrimental effect on generalization with respect to sequence length.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d3081a6-de7d-4e3c-9f67-89313c8a087aCited by top-tier papers1
Ask how each one uses itBuilds on4
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsKaiyue Wen, Yuchen Li, Bingbin Liu, Andrej RisteskiNeurIPS 2023 · 32 citations
- Self-Attention Networks Can Process Bounded Hierarchical LanguagesShunyu Yao, Binghui Peng, Christos H. Papadimitriou, Karthik NarasimhanACL 2021
Related papers
- A Formal Framework for Understanding Length Generalization in TransformersXinting Huang, Andy Yang, Satwik Bhattamishra, Yash Raj Sarrof et al.ICLR 2025
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- How Transformers Learn Structured Data: Insights From Hierarchical FilteringJerome Garnier-Brun, Marc Mézard, Emanuele Moscato, Luca SagliettiICML 2025
- SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingKrishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh et al.EMNLP 2025
- Provable Memorization Capacity of TransformersJunghwan Kim, Michelle Kim, Barzan MozafariICLR 2023
