Algebraic Positional Encodings
Konstantinos Kogkalidis, Jean-Philippe Bernardy, Vikas Garg
Abstract
We introduce a novel positional encoding strategy for Transformer-style models, addressing the shortcomings of existing, often ad hoc, approaches. Our framework provides a flexible mapping from the algebraic specification of a domain to an interpretation as orthogonal operators. This design preserves the algebraic characteristics of the source domain, ensuring that the model upholds its desired structural properties. Our scheme can accommodate various structures, ncluding sequences, grids and trees, as well as their compositions. We conduct a series of experiments to demonstrate the practical applicability of our approach. Results suggest performance on par with or surpassing the current state-of-the-art, without hyper-parameter optimizations or"task search"of any kind. Code is available at https://github.com/konstantinosKokos/ape.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d30cb33-b813-41fd-af17-76c9bd41d787Cited by top-tier papers5
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan et al.NeurIPS 2025 · 36 citations
- Learning Structure-Aware Representations of Dependent TypesKonstantinos Kogkalidis, Orestis Melkonian, Jean-Philippe BernardyNeurIPS 2024 · 6 citations
- Group Representational Position EncodingYifan Zhang, Zixiang Chen, Yifeng Liu, Zhen Qin et al.ICLR 2026 · 5 citations
- FiX: Introducing Fine-grained Forget Gate into Softmax AttentionRunzhong Li, Renjie Liu, Qing Li, Bo TangICML 2026
- Positional Encoding meets Persistent Homology on GraphsYogesh Verma, Amauri H. Souza, Vikas K. GargICML 2025
Builds on5
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- SOFT: Softmax-free Transformer with Linear ComplexityJiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu et al.NeurIPS 2021 · 232 citations
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 193 citations
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li et al.ICLR 2020 · 134 citations
- Learning Structure-Aware Representations of Dependent TypesKonstantinos Kogkalidis, Orestis Melkonian, Jean-Philippe BernardyNeurIPS 2024 · 6 citations
Related papers
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential FeaturesKaichen Xu, Yihang Du, Mianpeng Liu, Zimu Yu et al.NeurIPS 2025 · 4 citations
- Decoupling The "What" and "Where" With Polar Coordinate Positional EmbeddingAnand Gopalakrishnan, Róbert Csordás, Jürgen Schmidhuber, Michael MozerICML 2026 · 7 citations
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang et al.NeurIPS 2024 · 42 citations
- Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional EncodingDaichi Hayakawa, Issei SatoACL 2025
- Transformers over Directed Acyclic GraphsYuankai Luo, Veronika Thost, Lei ShiNeurIPS 2023 · 43 citations
