Thin Phase 2 preview
arXiv:1706.03762
cs.CL
cs.LG

Paper Ingestion Preview

Feed a research paper into Arbitro, extract facts and theories, then expose the support/dependency graph for review.

View paper record
Upload test

Upload a research paper

Select a PDF to exercise the paper-ingestion path. In this preview, the uploaded file is matched to the static Attention Is All You Need extraction so the UI flow can be tested before backend PDF parsing is wired up.

The selected PDF is held in the browser file control; the static extraction below shows the expected parsed output for this test paper.

Source artifact

Attention Is All You Need

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks. The Transformer is a new simple network architecture based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

Authors

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit et al.

Published

June 12, 2017

Taxonomy placement

Formal sciencesComputer science / machine learningSequence transduction with attention mechanismsPrimary source artifact
Extraction pipeline

Static fixture status

Metadata fetched
Preview
Text extracted
Preview
Claims extracted
Preview
Relationships inferred
Preview

This is intentionally not live recursive ingestion yet. It pulls forward just enough Phase 2 taxonomy/evidence behavior to make the MVP product loop visible.

Extracted facts & theories

These cards reuse the existing fact UI, but the data is now tied to a specific primary source artifact.

THEORY
PENDING
Jun 4, 2026

The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.

Confidence92%
Consensus88%
#arxiv
#transformer
#attention
#architecture
#paper:1706.03762
Sources:
arxiv.org+1 more sources
THEORY
PENDING
Jun 4, 2026

Multi-head attention lets the model attend to information from different representation subspaces at different positions.

Confidence89%
Consensus82%
#multi-head-attention
#representation
#attention
#paper:1706.03762
Sources:
arxiv.org+1 more sources
FACT
PENDING
Jun 4, 2026

Because the Transformer has no recurrence or convolution, positional encodings are added so the model can use token order.

Confidence91%
Consensus86%
#positional-encoding
#sequence-order
#transformer
#paper:1706.03762
Sources:
arxiv.org+1 more sources
FACT
PENDING
Jun 4, 2026

Scaled dot-product attention computes attention weights from queries and keys, then applies them to values.

Confidence93%
Consensus87%
#scaled-dot-product-attention
#queries
#keys
#values
#paper:1706.03762
Sources:
arxiv.org+1 more sources
HYPOTHESIS
PENDING
Jun 4, 2026

Self-attention reduces the amount of sequential computation compared with recurrent sequence models.

Confidence86%
Consensus80%
#self-attention
#parallelism
#recurrent-networks
#paper:1706.03762
Sources:
arxiv.org+1 more sources
FACT
PENDING
Jun 4, 2026

The Transformer achieved strong machine-translation results on WMT 2014 English-German and English-French benchmarks.

Confidence88%
Consensus83%
#machine-translation
#wmt-2014
#benchmark
#paper:1706.03762
Sources:
arxiv.org+1 more sources

References/Citations

40 extracted
2016
1 in-text mention
layernorm2016

Jimmy Lei Ba et al. (2016). Layer normalization. arXiv preprint arXiv:1607.06450.

Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E Hinton

Raw reference text

Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

2014
5 in-text mentions
bahdanau2014neural

Dzmitry Bahdanau et al. (2014). Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473.

Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio

Raw reference text

Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473, 2014

2017
2 in-text mentions
DBLP:journals/corr/BritzGLL17

Denny Britz et al. (2017). Massive exploration of neural machine translation architectures. CoRR, abs/1703.03906.

Denny Britz, Anna Goldie, Minh-Thang Luong, Quoc V. Le

Raw reference text

Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc V. Le. Massive exploration of neural machine translation architectures. CoRR, abs/1703.03906, 2017

2016
1 in-text mention
cheng2016long

Jianpeng Cheng et al. (2016). Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733.

Jianpeng Cheng, Li Dong, Mirella Lapata

Raw reference text

Jianpeng Cheng, Li Dong, and Mirella Lapata. Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733, 2016

2014
2 in-text mentions
cho2014learning

Kyunghyun Cho et al. (2014). Learning phrase representations using rnn encoder-decoder for statistical machine translation. CoRR, abs/1406.1078.

Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, Yoshua Bengio

Raw reference text

Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014

2016
1 in-text mention
xception2016

Francois Chollet (2016). Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357.

Francois Chollet

Raw reference text

Francois Chollet. Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357, 2016

2014
1 in-text mention
gruEval14

Junyoung Chung et al. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555.

Junyoung Chung, Caglar Gülcehre, Kyunghyun Cho, Yoshua Bengio

Raw reference text

Junyoung Chung, Caglar Gülcehre, Kyunghyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555, 2014

2016
3 in-text mentions
dyer-rnng:16

Chris Dyer et al. (2016). Recurrent neural network grammars. In Proc. of NAACL.

Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith

Raw reference text

Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. Recurrent neural network grammars. In Proc. of NAACL, 2016

2017
8 in-text mentions
JonasFaceNet2017

Jonas Gehring et al. (2017). Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122v2.

Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin

Raw reference text

Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122v2, 2017

2013
1 in-text mention
graves2013generating

Alex Graves (2013). Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850.

Alex Graves

Raw reference text

Alex Graves. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013

2016
1 in-text mention
he2016deep

Kaiming He et al. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778.

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

Raw reference text

Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778, 2016

2001
2 in-text mentions
hochreiter2001gradient

Sepp Hochreiter et al. (2001). Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001.

Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber

Raw reference text

Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001

1997
1 in-text mention
hochreiter1997

Sepp Hochreiter and Jürgen Schmidhuber (1997). Long short-term memory. Neural computation, 9(8):1735--1780.

Sepp Hochreiter, Jürgen Schmidhuber

Raw reference text

Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735--1780, 1997

2009
1 in-text mention
huang-harper:2009:EMNLP

Zhongqiang Huang and Mary Harper (2009). Self-training PCFG grammars with latent annotations across languages. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 832--841. ACL, August.

Zhongqiang Huang, Mary Harper

Raw reference text

Zhongqiang Huang and Mary Harper. Self-training PCFG grammars with latent annotations across languages. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 832--841. ACL, August 2009

2016
1 in-text mention
jozefowicz2016exploring

Rafal Jozefowicz et al. (2016). Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.

Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, Yonghui Wu

Raw reference text

Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016

2016
1 in-text mention
extendedngpu

ukasz Kaiser and Samy Bengio (2016). Can active memory replace attention?. In Advances in Neural Information Processing Systems, (NIPS).

ukasz Kaiser, Samy Bengio

Raw reference text

ukasz Kaiser and Samy Bengio. Can active memory replace attention? In Advances in Neural Information Processing Systems, (NIPS), 2016

2016
1 in-text mention
neural_gpu

ukasz Kaiser and Ilya Sutskever (2016). Neural GPUs learn algorithms. In International Conference on Learning Representations (ICLR).

ukasz Kaiser, Ilya Sutskever

Raw reference text

ukasz Kaiser and Ilya Sutskever. Neural GPUs learn algorithms. In International Conference on Learning Representations (ICLR), 2016

2017
4 in-text mentions
NalBytenet2017

Nal Kalchbrenner et al. (2017). Neural machine translation in linear time. arXiv preprint arXiv:1610.10099v2.

Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, Koray Kavukcuoglu

Raw reference text

Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. Neural machine translation in linear time. arXiv preprint arXiv:1610.10099v2, 2017

2017
1 in-text mention
structuredAttentionNetworks

Yoon Kim et al. (2017). Structured attention networks. In International Conference on Learning Representations.

Yoon Kim, Carl Denton, Luong Hoang, Alexander M. Rush

Raw reference text

Yoon Kim, Carl Denton, Luong Hoang, and Alexander M. Rush. Structured attention networks. In International Conference on Learning Representations, 2017

2015
1 in-text mention
kingma2014adam

Diederik Kingma and Jimmy Ba (2015). Adam: A method for stochastic optimization. In ICLR.

Diederik Kingma, Jimmy Ba

Raw reference text

Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015

2017
1 in-text mention
Kuchaiev2017Factorization

Oleksii Kuchaiev and Boris Ginsburg (2017). Factorization tricks for LSTM networks. arXiv preprint arXiv:1703.10722.

Oleksii Kuchaiev, Boris Ginsburg

Raw reference text

Oleksii Kuchaiev and Boris Ginsburg. Factorization tricks for LSTM networks. arXiv preprint arXiv:1703.10722, 2017

2017
1 in-text mention
lin2017structured

Zhouhan Lin et al. (2017). A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130.

Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, Yoshua Bengio

Raw reference text

Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130, 2017

2015
1 in-text mention
multiseq2seq

Minh-Thang Luong et al. (2015). Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114.

Minh-Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, Lukasz Kaiser

Raw reference text

Minh-Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114, 2015

2015
1 in-text mention
luong2015effective

Minh-Thang Luong et al. (2015). Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025.

Minh-Thang Luong, Hieu Pham, Christopher D Manning

Raw reference text

Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025, 2015

1993
1 in-text mention
marcus1993building

Mitchell P Marcus et al. (1993). Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313--330.

Mitchell P Marcus, Mary Ann Marcinkiewicz, Beatrice Santorini

Raw reference text

Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313--330, 1993

2006
1 in-text mention
mcclosky-etAl:2006:NAACL

David McClosky et al. (2006). Effective self-training for parsing. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152--159. ACL, June.

David McClosky, Eugene Charniak, Mark Johnson

Raw reference text

David McClosky, Eugene Charniak, and Mark Johnson. Effective self-training for parsing. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152--159. ACL, June 2006

2016
2 in-text mentions
decomposableAttnModel

Ankur Parikh et al. (2016). A decomposable attention model. In Empirical Methods in Natural Language Processing.

Ankur Parikh, Oscar Täckström, Dipanjan Das, Jakob Uszkoreit

Raw reference text

Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. A decomposable attention model. In Empirical Methods in Natural Language Processing, 2016

2017
1 in-text mention
paulus2017deep

Romain Paulus et al. (2017). A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304.

Romain Paulus, Caiming Xiong, Richard Socher

Raw reference text

Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017

2006
2 in-text mentions
petrov-EtAl:2006:ACL

Slav Petrov et al. (2006). Learning accurate, compact, and interpretable tree annotation. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the ACL, pages 433--440. ACL, July.

Slav Petrov, Leon Barrett, Romain Thibaux, Dan Klein

Raw reference text

Slav Petrov, Leon Barrett, Romain Thibaux, and Dan Klein. Learning accurate, compact, and interpretable tree annotation. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the ACL, pages 433--440. ACL, July 2006

2016
1 in-text mention
press2016using

Ofir Press and Lior Wolf (2016). Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859.

Ofir Press, Lior Wolf

Raw reference text

Ofir Press and Lior Wolf. Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859, 2016

2015
1 in-text mention
sennrich2015neural

Rico Sennrich et al. (2015). Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909.

Rico Sennrich, Barry Haddow, Alexandra Birch

Raw reference text

Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015

2017
2 in-text mentions
shazeer2017outrageously

Noam Shazeer et al. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.

Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean

Raw reference text

Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017

1929
1 in-text mention
srivastava2014dropout

Nitish Srivastava et al. (1929). Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):.

Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, Ruslan Salakhutdinov

Raw reference text

Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929--1958, 2014

2015
1 in-text mention
sukhbaatar2015

Sainbayar Sukhbaatar et al. (2015). End-to-end memory networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 2440--2448. Curran Associates, Inc.

Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus

Raw reference text

Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. End-to-end memory networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 2440--2448. Curran Associates, Inc., 2015

2014
2 in-text mentions
sutskever14

Ilya Sutskever et al. (2014). Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104--3112.

Ilya Sutskever, Oriol Vinyals, Quoc VV Le

Raw reference text

Ilya Sutskever, Oriol Vinyals, and Quoc VV Le. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104--3112, 2014

2015
1 in-text mention
DBLP:journals/corr/SzegedyVISW15

Christian Szegedy et al. (2015). Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567.

Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, Zbigniew Wojna

Raw reference text

Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015

2015
5 in-text mentions
KVparse15

Vinyals & Kaiser et al. (2015). Grammar as a foreign language. In Advances in Neural Information Processing Systems.

Vinyals & Kaiser, Koo, Petrov, Sutskever, Hinton

Raw reference text

Vinyals & Kaiser, Koo, Petrov, Sutskever, and Hinton. Grammar as a foreign language. In Advances in Neural Information Processing Systems, 2015

2016
8 in-text mentions
wu2016google

Yonghui Wu et al. (2016). Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.

Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey

Raw reference text

Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016

2016
2 in-text mentions
DBLP:journals/corr/ZhouCWLX16

Jie Zhou et al. (2016). Deep recurrent models with fast-forward connections for neural machine translation. CoRR, abs/1606.04199.

Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, Wei Xu

Raw reference text

Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. Deep recurrent models with fast-forward connections for neural machine translation. CoRR, abs/1606.04199, 2016

2013
2 in-text mentions
zhu-EtAl:2013:ACL

Muhua Zhu et al. (2013). Fast and accurate shift-reduce constituent parsing. In Proceedings of the 51st Annual Meeting of the ACL (Volume 1: Long Papers), pages 434--443. ACL, August.

Muhua Zhu, Yue Zhang, Wenliang Chen, Min Zhang, Jingbo Zhu

Raw reference text

Muhua Zhu, Yue Zhang, Wenliang Chen, Min Zhang, and Jingbo Zhu. Fast and accurate shift-reduce constituent parsing. In Proceedings of the 51st Annual Meeting of the ACL (Volume 1: Long Papers), pages 434--443. ACL, August 2013. thebibliography

Relationships

The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.
depends on
Scaled dot-product attention computes attention weights from queries and keys, then applies them to values.

The attention-only architecture depends on an attention primitive; the paper instantiates that primitive as scaled dot-product attention over queries, keys, and values.

The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.
extends
Multi-head attention lets the model attend to information from different representation subspaces at different positions.

Multi-head attention extends the base attention mechanism by running attention in multiple learned representation subspaces.

The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.
depends on
Because the Transformer has no recurrence or convolution, positional encodings are added so the model can use token order.

Removing recurrence and convolution removes built-in order bias, so the architecture depends on positional encodings to represent sequence order.

Self-attention reduces the amount of sequential computation compared with recurrent sequence models.
supports
The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.

The claim about reduced sequential computation supports the paper's argument for replacing recurrence with self-attention.

The Transformer achieved strong machine-translation results on WMT 2014 English-German and English-French benchmarks.
is evidence for
The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.

The WMT benchmark results are empirical evidence that the attention-only architecture can perform well on sequence transduction.