Paper Ingestion Preview
Feed a research paper into Arbitro, extract facts and theories, then expose the support/dependency graph for review.
Upload a research paper
Select a PDF to exercise the paper-ingestion path. In this preview, the uploaded file is matched to the static Attention Is All You Need extraction so the UI flow can be tested before backend PDF parsing is wired up.
The selected PDF is held in the browser file control; the static extraction below shows the expected parsed output for this test paper.
Attention Is All You Need
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks. The Transformer is a new simple network architecture based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
Authors
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit et al.
Published
June 12, 2017
Taxonomy placement
Static fixture status
This is intentionally not live recursive ingestion yet. It pulls forward just enough Phase 2 taxonomy/evidence behavior to make the MVP product loop visible.
Extracted facts & theories
These cards reuse the existing fact UI, but the data is now tied to a specific primary source artifact.
The Transformer architecture relies entirely on attention mechanisms and avoids recurrent and convolutional layers for sequence transduction.
Multi-head attention lets the model attend to information from different representation subspaces at different positions.
Because the Transformer has no recurrence or convolution, positional encodings are added so the model can use token order.
Scaled dot-product attention computes attention weights from queries and keys, then applies them to values.
Self-attention reduces the amount of sequential computation compared with recurrent sequence models.
The Transformer achieved strong machine-translation results on WMT 2014 English-German and English-French benchmarks.
References/Citations
Dzmitry Bahdanau et al. (2014). Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473.
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio
Raw reference text
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473, 2014
Denny Britz et al. (2017). Massive exploration of neural machine translation architectures. CoRR, abs/1703.03906.
Denny Britz, Anna Goldie, Minh-Thang Luong, Quoc V. Le
Raw reference text
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc V. Le. Massive exploration of neural machine translation architectures. CoRR, abs/1703.03906, 2017
Jianpeng Cheng et al. (2016). Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733.
Jianpeng Cheng, Li Dong, Mirella Lapata
Raw reference text
Jianpeng Cheng, Li Dong, and Mirella Lapata. Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733, 2016
Kyunghyun Cho et al. (2014). Learning phrase representations using rnn encoder-decoder for statistical machine translation. CoRR, abs/1406.1078.
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, Yoshua Bengio
Raw reference text
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014
Francois Chollet (2016). Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357.
Francois Chollet
Raw reference text
Francois Chollet. Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357, 2016
Junyoung Chung et al. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555.
Junyoung Chung, Caglar Gülcehre, Kyunghyun Cho, Yoshua Bengio
Raw reference text
Junyoung Chung, Caglar Gülcehre, Kyunghyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555, 2014
Chris Dyer et al. (2016). Recurrent neural network grammars. In Proc. of NAACL.
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith
Raw reference text
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. Recurrent neural network grammars. In Proc. of NAACL, 2016
Jonas Gehring et al. (2017). Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122v2.
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin
Raw reference text
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122v2, 2017
Kaiming He et al. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun
Raw reference text
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770--778, 2016
Sepp Hochreiter et al. (2001). Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001.
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber
Raw reference text
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Sepp Hochreiter and Jürgen Schmidhuber (1997). Long short-term memory. Neural computation, 9(8):1735--1780.
Sepp Hochreiter, Jürgen Schmidhuber
Raw reference text
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735--1780, 1997
Zhongqiang Huang and Mary Harper (2009). Self-training PCFG grammars with latent annotations across languages. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 832--841. ACL, August.
Zhongqiang Huang, Mary Harper
Raw reference text
Zhongqiang Huang and Mary Harper. Self-training PCFG grammars with latent annotations across languages. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 832--841. ACL, August 2009
Rafal Jozefowicz et al. (2016). Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, Yonghui Wu
Raw reference text
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016
ukasz Kaiser and Samy Bengio (2016). Can active memory replace attention?. In Advances in Neural Information Processing Systems, (NIPS).
ukasz Kaiser, Samy Bengio
Raw reference text
ukasz Kaiser and Samy Bengio. Can active memory replace attention? In Advances in Neural Information Processing Systems, (NIPS), 2016
ukasz Kaiser and Ilya Sutskever (2016). Neural GPUs learn algorithms. In International Conference on Learning Representations (ICLR).
ukasz Kaiser, Ilya Sutskever
Raw reference text
ukasz Kaiser and Ilya Sutskever. Neural GPUs learn algorithms. In International Conference on Learning Representations (ICLR), 2016
Nal Kalchbrenner et al. (2017). Neural machine translation in linear time. arXiv preprint arXiv:1610.10099v2.
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, Koray Kavukcuoglu
Raw reference text
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. Neural machine translation in linear time. arXiv preprint arXiv:1610.10099v2, 2017
Yoon Kim et al. (2017). Structured attention networks. In International Conference on Learning Representations.
Yoon Kim, Carl Denton, Luong Hoang, Alexander M. Rush
Raw reference text
Yoon Kim, Carl Denton, Luong Hoang, and Alexander M. Rush. Structured attention networks. In International Conference on Learning Representations, 2017
Oleksii Kuchaiev and Boris Ginsburg (2017). Factorization tricks for LSTM networks. arXiv preprint arXiv:1703.10722.
Oleksii Kuchaiev, Boris Ginsburg
Raw reference text
Oleksii Kuchaiev and Boris Ginsburg. Factorization tricks for LSTM networks. arXiv preprint arXiv:1703.10722, 2017
Zhouhan Lin et al. (2017). A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130.
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, Yoshua Bengio
Raw reference text
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130, 2017
Minh-Thang Luong et al. (2015). Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114.
Minh-Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, Lukasz Kaiser
Raw reference text
Minh-Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114, 2015
Minh-Thang Luong et al. (2015). Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025.
Minh-Thang Luong, Hieu Pham, Christopher D Manning
Raw reference text
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025, 2015
Mitchell P Marcus et al. (1993). Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313--330.
Mitchell P Marcus, Mary Ann Marcinkiewicz, Beatrice Santorini
Raw reference text
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313--330, 1993
David McClosky et al. (2006). Effective self-training for parsing. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152--159. ACL, June.
David McClosky, Eugene Charniak, Mark Johnson
Raw reference text
David McClosky, Eugene Charniak, and Mark Johnson. Effective self-training for parsing. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152--159. ACL, June 2006
Ankur Parikh et al. (2016). A decomposable attention model. In Empirical Methods in Natural Language Processing.
Ankur Parikh, Oscar Täckström, Dipanjan Das, Jakob Uszkoreit
Raw reference text
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. A decomposable attention model. In Empirical Methods in Natural Language Processing, 2016
Romain Paulus et al. (2017). A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304.
Romain Paulus, Caiming Xiong, Richard Socher
Raw reference text
Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017
Slav Petrov et al. (2006). Learning accurate, compact, and interpretable tree annotation. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the ACL, pages 433--440. ACL, July.
Slav Petrov, Leon Barrett, Romain Thibaux, Dan Klein
Raw reference text
Slav Petrov, Leon Barrett, Romain Thibaux, and Dan Klein. Learning accurate, compact, and interpretable tree annotation. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the ACL, pages 433--440. ACL, July 2006
Ofir Press and Lior Wolf (2016). Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859.
Ofir Press, Lior Wolf
Raw reference text
Ofir Press and Lior Wolf. Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859, 2016
Rico Sennrich et al. (2015). Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909.
Rico Sennrich, Barry Haddow, Alexandra Birch
Raw reference text
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015
Noam Shazeer et al. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean
Raw reference text
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017
Nitish Srivastava et al. (1929). Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):.
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, Ruslan Salakhutdinov
Raw reference text
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929--1958, 2014
Sainbayar Sukhbaatar et al. (2015). End-to-end memory networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 2440--2448. Curran Associates, Inc.
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus
Raw reference text
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. End-to-end memory networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 2440--2448. Curran Associates, Inc., 2015
Ilya Sutskever et al. (2014). Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104--3112.
Ilya Sutskever, Oriol Vinyals, Quoc VV Le
Raw reference text
Ilya Sutskever, Oriol Vinyals, and Quoc VV Le. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104--3112, 2014
Christian Szegedy et al. (2015). Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567.
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, Zbigniew Wojna
Raw reference text
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015
Vinyals & Kaiser et al. (2015). Grammar as a foreign language. In Advances in Neural Information Processing Systems.
Vinyals & Kaiser, Koo, Petrov, Sutskever, Hinton
Raw reference text
Vinyals & Kaiser, Koo, Petrov, Sutskever, and Hinton. Grammar as a foreign language. In Advances in Neural Information Processing Systems, 2015
Yonghui Wu et al. (2016). Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey
Raw reference text
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016
Jie Zhou et al. (2016). Deep recurrent models with fast-forward connections for neural machine translation. CoRR, abs/1606.04199.
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, Wei Xu
Raw reference text
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. Deep recurrent models with fast-forward connections for neural machine translation. CoRR, abs/1606.04199, 2016
Muhua Zhu et al. (2013). Fast and accurate shift-reduce constituent parsing. In Proceedings of the 51st Annual Meeting of the ACL (Volume 1: Long Papers), pages 434--443. ACL, August.
Muhua Zhu, Yue Zhang, Wenliang Chen, Min Zhang, Jingbo Zhu
Raw reference text
Muhua Zhu, Yue Zhang, Wenliang Chen, Min Zhang, and Jingbo Zhu. Fast and accurate shift-reduce constituent parsing. In Proceedings of the 51st Annual Meeting of the ACL (Volume 1: Long Papers), pages 434--443. ACL, August 2013. thebibliography
Relationships
The attention-only architecture depends on an attention primitive; the paper instantiates that primitive as scaled dot-product attention over queries, keys, and values.
Multi-head attention extends the base attention mechanism by running attention in multiple learned representation subspaces.
Removing recurrence and convolution removes built-in order bias, so the architecture depends on positional encodings to represent sequence order.
The claim about reduced sequential computation supports the paper's argument for replacing recurrence with self-attention.
The WMT benchmark results are empirical evidence that the attention-only architecture can perform well on sequence transduction.