← Research papers
2025unread

Attention Is All You Need

Ashish VaswaniNoam ShazeerNiki ParmarJakob UszkoreitLlion JonesAidan N.GomezLukasz KaiserIllia Polosukhin
Publisher page
Open graph

Citations

0

Open access

No

Source

crossref

OpenAlex

Not enriched

DOI

10.65215/nxvz2v36

Abstract

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.

Collections

Add to collection

Paper intelligence

Research analysis

Confidence 95%

16 source chunks

Summary

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU.

Plain-language summary

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU.

Research problem

We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. ∗Equal contribution. The fundamental constraint of sequential computation, however, remains.

Methodology

We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. Our model achieves 28.4 BLEU on the WMT 2014 English- to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU.

Main findings

Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. Our model achieves 28.4 BLEU on the WMT 2014 English- to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research. †Work performed while at Google Brain. ‡Work performed while at Google Research. Recent work has achieved significant improvements in computational efficiency through factorization tricks [21] and conditional computation [32], while also improving model performance in case of the latter. While for small values of dk the two mechanisms perform similarly, additive attention outperforms dot product attention without scaling for larger values of dk [3].

Key contributions

Contribution 1

We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

Contribution 2

We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. ∗Equal contribution.

Contribution 3

Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work.

Contribution 4

Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations.

Contribution 5

In this work we propose the Transformer, a model architecture eschewing recurrence and instead relying entirely on an attention mechanism to draw global dependencies between input and output.

Contribution 6

To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence- aligned RNNs or convolution. -independent sentence representations [4, 27, 28, 22].

Contribution 7

To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence- aligned RNNs or convolution.

Limitations

  • A single convolutional layer with kernel width k < n does not connect all pairs of input and output positions.

Future work

  • We plan to investigate this approach further in future work.
  • We plan to extend the Transformer to problems involving input and output modalities other than text and to investigate local, restricted attention mechanisms to efficiently handle large inputs and outputs such as images, audio and video.

Models and methods

TransformerLSTM

Datasets

English-German

5.1 Training Data and Batching We trained on the standard WMT 2014 English-German dataset consisting of about 4.5 million sentence pairs.

English-French

For English-French, we used the significantly larger WMT 2014 English-French dataset consisting of 36M sentences and split tokens into a 32000 word-piece vocabulary [38].

Dataset 3

Building a large annotated corpus of english: The penn treebank.

Evaluation metrics

accuracyprecisionF1BLEUperplexity

Keywords

machine learningdeep learningcomputer visionnatural language processing

No graph connections yet.

Sync citations or add papers to shared collections to build this network.

Knowledge graph

Citation network

Explore references, papers that cite this work and related papers in your Codex library.

References

0

No references have been linked yet.

Cited by

0

No saved paper is currently linked as citing this work.

Related papers

0

Add papers to shared collections or enrich their topics to find related work.

Research workspace

Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.

Attach PDF

Upload the research paper so Codex can extract, chunk and search its contents.

Paper resources

Attention Is All You Need.pdf

Status: completed16 chunks40091 charactersapplication/pdf