- Type
- paper
- Year
- 2017
- By
- Vaswani et al. (Google)
- Publisher
- Advances in Neural Information Processing Systems (NeurIPS)
- DOI
- 10.48550/arXiv.1706.03762
- External
- arxiv.org/abs/1706.03762
"Attention Is All You Need" presents the Transformer architecture, a neural network design that relies entirely on attention mechanisms rather than recurrence or convolution. The paper introduces multi-head self-attention, allowing models to attend to different representation subspaces simultaneously, and demonstrates that this approach achieves state-of-the-art results on machine translation tasks while being more parallelizable and requiring less training time than previous sequence-to-sequence models.
The architecture consists of an encoder-decoder structure with stacked layers of multi-head attention and feed-forward networks. The paper includes positional encodings to inject sequence order information, since the model has no inherent notion of token position. Extensive experiments on WMT 2014 English-to-German and English-to-French translation benchmarks show superior performance compared to existing approaches.
This work has become foundational to modern deep learning. The Transformer architecture forms the basis of BERT, GPT, and virtually all contemporary large language models. It has been adapted for vision, audio, and multimodal tasks, making it one of the most influential papers in machine learning history, with over 100,000 citations.
Last updated 31 August 2026