paper

Attention Is All You Need

Foundational 2017 paper introducing the Transformer architecture, which replaced recurrence with self-attention and became the basis for modern large language models.

Attention Is All You Need
Type
paper
Year
2017
By
Vaswani et al. (Google)
Publisher
Advances in Neural Information Processing Systems (NeurIPS)
DOI
10.48550/arXiv.1706.03762

"Attention Is All You Need" presents the Transformer architecture, a neural network design that relies entirely on attention mechanisms rather than recurrence or convolution. The paper introduces multi-head self-attention, allowing models to attend to different representation subspaces simultaneously, and demonstrates that this approach achieves state-of-the-art results on machine translation tasks while being more parallelizable and requiring less training time than previous sequence-to-sequence models.

The architecture consists of an encoder-decoder structure with stacked layers of multi-head attention and feed-forward networks. The paper includes positional encodings to inject sequence order information, since the model has no inherent notion of token position. Extensive experiments on WMT 2014 English-to-German and English-to-French translation benchmarks show superior performance compared to existing approaches.

This work has become foundational to modern deep learning. The Transformer architecture forms the basis of BERT, GPT, and virtually all contemporary large language models. It has been adapted for vision, audio, and multimodal tasks, making it one of the most influential papers in machine learning history, with over 100,000 citations.

Last updated 31 August 2026