Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:General-purpose, long-context autoregressive modeling with Perceiver AR

Feb 15, 2022

Curtis Hawthorne, Andrew Jaegle, Cătălina Cangea, Sebastian Borgeaud, Charlie Nash, Mateusz Malinowski, Sander Dieleman, Oriol Vinyals, Matthew Botvinick, Ian Simon(+5 more)

Figure 1 for General-purpose, long-context autoregressive modeling with Perceiver AR

Figure 2 for General-purpose, long-context autoregressive modeling with Perceiver AR

Figure 3 for General-purpose, long-context autoregressive modeling with Perceiver AR

Figure 4 for General-purpose, long-context autoregressive modeling with Perceiver AR

Share this with someone who'll enjoy it:

Abstract:Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively expensive to scale to the number of inputs and layers needed to capture this long-range structure. We develop Perceiver AR, an autoregressive, modality-agnostic architecture which uses cross-attention to map long-range inputs to a small number of latents while also maintaining end-to-end causal masking. Perceiver AR can directly attend to over a hundred thousand tokens, enabling practical long-context density estimation without the need for hand-crafted sparsity patterns or memory mechanisms. When trained on images or music, Perceiver AR generates outputs with clear long-term coherence and structure. Our architecture also obtains state-of-the-art likelihood on long-sequence benchmarks, including 64 x 64 ImageNet images and PG-19 books.

View paper on

Share this with someone who'll enjoy it:

Title:General-purpose, long-context autoregressive modeling with Perceiver AR

Paper and Code