A Simple and Efficient Tensor Calculus
Sören Laue, Matthias Mitterreiter, Joachim Giesen
Abstract
Computing derivatives of tensor expressions, also known as tensor calculus, is a fundamental task in machine learning. A key concern is the efficiency of evaluating the expressions and their derivatives that hinges on the representation of these expressions. Recently, an algorithm for computing higher order derivatives of tensor expressions like Jacobians or Hessians has been introduced that is a few orders of magnitude faster than previous state-of-the-art approaches. Unfortunately, the approach is based on Ricci notation and hence cannot be incorporated into automatic differentiation frameworks from deep learning like TensorFlow, PyTorch, autograd, or JAX that use the simpler Einstein notation. This leaves two options, to either change the underlying tensor representation in these frameworks or to develop a new, provably correct algorithm based on Einstein notation. Obviously, the first option is impractical. Hence, we pursue the second option. Here, we show that using Ricci notation is not necessary for an efficient tensor calculus and develop an equally efficient method for the simpler Einstein notation. It turns out that turning to Einstein notation enables further improvements that lead to even better efficiency. The methods that are described in this paper have been implemented in the online tool www.MatrixCalculus.org for computing derivatives of matrix and tensor expressions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1f641da-d48f-4081-9332-37c02b90cf83Cited by top-tier papers8
- Tensor Relational Algebra for Distributed Machine Learning System DesignBinhang Yuan, Dimitrije Jankov, Jia Zou, Yuxin Tang et al.VLDB 2021 · 33 citations
- Why Capsule Neural Networks Do Not Scale: Challenging the Dynamic Parse-Tree AssumptionMatthias Mitterreiter, Marcel Koch, Joachim Giesen, Sören LaueAAAI 2023 · 17 citations
- Optimization for Classical Machine Learning Problems on the GPUSören Laue, Mark Blacher, Joachim GiesenAAAI 2022 · 9 citations
- Modeling the Impact of Timeline Algorithms on Opinion Dynamics Using Low-rank UpdatesTianyi Zhou, Stefan Neumann, Kiran Garimella, Aristides GionisWWW 2024 · 7 citations
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 5 citations
Builds on1
Related papers
- Einops: Clear and Reliable Tensor Manipulations with Einstein-like NotationAlex RogozhnikovICLR 2022 · 124 citations
- EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel ExecutionDaniel Bourgeois, Zhimin Ding, Dimitrije Jankov, Jiehui Li et al.VLDB 2025 · 4 citations
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 114 citations
- It's All Just Vectorization: einx, a Universal Notation for Tensor OperationsFlorian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael ArensICLR 2026 · 1 citation
- EINNET: Optimizing Tensor Programs with Derivation-Based TransformationsLiyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu et al.OSDI 2023 · 19 citations
