Algebraic Positional Encodings
Konstantinos Kogkalidis, Jean-Philippe Bernardy, Vikas Garg
摘要
We introduce a novel positional encoding strategy for Transformer-style models, addressing the shortcomings of existing, often ad hoc, approaches. Our framework provides a flexible mapping from the algebraic specification of a domain to an interpretation as orthogonal operators. This design preserves the algebraic characteristics of the source domain, ensuring that the model upholds its desired structural properties. Our scheme can accommodate various structures, ncluding sequences, grids and trees, as well as their compositions. We conduct a series of experiments to demonstrate the practical applicability of our approach. Results suggest performance on par with or surpassing the current state-of-the-art, without hyper-parameter optimizations or"task search"of any kind. Code is available at https://github.com/konstantinosKokos/ape.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan 等NeurIPS 2025 · 被引用 36 次
- Learning Structure-Aware Representations of Dependent TypesKonstantinos Kogkalidis, Orestis Melkonian, Jean-Philippe BernardyNeurIPS 2024 · 被引用 6 次
- Group Representational Position EncodingYifan Zhang, Zixiang Chen, Yifeng Liu, Zhen Qin 等ICLR 2026 · 被引用 5 次
- FiX: Introducing Fine-grained Forget Gate into Softmax AttentionRunzhong Li, Renjie Liu, Qing Li, Bo TangICML 2026
- Positional Encoding meets Persistent Homology on GraphsYogesh Verma, Amauri H. Souza, Vikas K. GargICML 2025
它引用的顶会 Paper5
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- SOFT: Softmax-free Transformer with Linear ComplexityJiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu 等NeurIPS 2021 · 被引用 232 次
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 被引用 193 次
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li 等ICLR 2020 · 被引用 134 次
- Learning Structure-Aware Representations of Dependent TypesKonstantinos Kogkalidis, Orestis Melkonian, Jean-Philippe BernardyNeurIPS 2024 · 被引用 6 次
相关 Paper
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential FeaturesKaichen Xu, Yihang Du, Mianpeng Liu, Zimu Yu 等NeurIPS 2025 · 被引用 4 次
- Decoupling The "What" and "Where" With Polar Coordinate Positional EmbeddingAnand Gopalakrishnan, Róbert Csordás, Jürgen Schmidhuber, Michael MozerICML 2026 · 被引用 7 次
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang 等NeurIPS 2024 · 被引用 42 次
- Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional EncodingDaichi Hayakawa, Issei SatoACL 2025
- Transformers over Directed Acyclic GraphsYuankai Luo, Veronika Thost, Lei ShiNeurIPS 2023 · 被引用 43 次
