Lune

ICLR2023Top-tier venue

Provable Memorization Capacity of Transformers

Junghwan Kim, Michelle Kim, Barzan Mozafari

2023Year
30Top-tier citations

Abstract

Quantifying memorization capacity is essential for understanding the expressiveness and generalizability of deep learning model architectures. However, the memorization capacity of the Transformer architecture has yet to be explored. In this work, we present the first study of the memorization capacity of the Transformer architecture. We prove that Transformers are capable of memorizing NN sequence-to-sequence mappings of length nn with dd-dimensional input tokens using O~(d+n+nN)\tilde{O}(d + n + \sqrt{nN}) parameters. Our theory supports memorization both with and without permutation equivariance, utilizing positional encodings in the latter case. Building on our theory, we also analyze the memorization capacity of Transformers in the sequence classification and language modeling tasks. To verify these theoretical findings, we conduct experiments analyzing the memorization capacity of Transformers in the natural language domain.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 46622d15-6c88-4b8a-9302-21035110d03a

Cited by top-tier papers30

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines