NS
Noam Shazeer
cs.LGcs.CLstat.MLcs.AIcs.NEcs.CVcs.SDcs.DCeess.ASeess.IV
On Valency
published · living versionsW_qeugbgvj·v1 · currentpublished
Preprints & journals
36 papers in the corpus · 2010–2022Faster Transformer Decoding: N-gram Masked Self-Attention2001.04589v2 · Ciprian Chelba, Mia Chen, Ankur Bapna et al.2020 · 12 citationsarXiv
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer1910.10683v4 · Colin Raffel, Noam Shazeer, Adam Roberts et al.2019 · 10,497 citationsarXiv
Attention Is All You Need1706.03762v7 · Ashish Vaswani, Noam Shazeer, Niki Parmar et al.2017 · 26,892 citationsarXivon Valency
PaLM: Scaling Language Modeling with Pathways2204.02311v5 · Aakanksha Chowdhery, Sharan Narang, Jacob Devlin et al.2022 · 2,138 citationsarXiv
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2101.03961v3 · William Fedus, Barret Zoph, Noam Shazeer2021 · 918 citationsarXiv
ST-MoE: Designing Stable and Transferable Sparse Expert Models2202.08906v2 · Barret Zoph, Irwan Bello, Sameer Kumar et al.2022 · 52 citationsarXiv
Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$2203.17189v1 · Adam Roberts, Hyung Won Chung, Anselm Levskaya et al.2022 · 47 citationsarXiv
LaMDA: Language Models for Dialog Applications2201.08239v3 · Romal Thoppilan, Daniel De Freitas, Jamie Hall et al.2022 · 696 citationsarXiv
Primer: Searching for Efficient Transformers for Language Modeling2109.08668v2 · David R. So, Wojciech Ma'nke, Hanxiao Liu et al.2021 · 11 citationsarXiv
GSPMD: General and Scalable Parallelization for ML Computation Graphs2105.04663v2 · Yuanzhong Xu, HyoukJoong Lee, Dehao Chen et al.2021 · 37 citationsarXiv
Do Transformer Modifications Transfer Across Implementations and Applications?2102.11972v2 · Sharan Narang, Hyung Won Chung, Yi Tay et al.2021 · 82 citationsarXiv
How Much Knowledge Can You Pack Into the Parameters of a Language Model?2002.08910v4 · Adam Roberts, Colin Raffel, Noam Shazeer2020 · 630 citationsarXiv
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding2006.16668v1 · Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu et al.2020 · 400 citationsarXiv
Talking-Heads Attention2003.02436v1 · Noam Shazeer, Zhenzhong Lan, Youlong Cheng et al.2020 · 49 citationsarXiv
GLU Variants Improve Transformer2002.05202v1 · Noam Shazeer2020 · 20 citationsarXiv
Fast Transformer Decoding: One Write-Head is All You Need1911.02150v1 · Noam Shazeer2019 · 26 citationsarXiv
High Resolution Medical Image Analysis with Spatial Partitioning1909.03108v3 · Le Hou, Youlong Cheng, Noam Shazeer et al.2019 · 28 citationsarXiv
Corpora Generation for Grammatical Error Correction1904.05780v1 · Jared Lichtarge, Chris Alberti, Shankar Kumar et al.2019 · 143 citationsarXiv
Music Transformer1809.04281v3 · Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit et al.2018 · 73 citationsarXiv
Blockwise Parallel Decoding for Deep Autoregressive Models1811.03115v1 · Mitchell Stern, Noam Shazeer, Jakob Uszkoreit2018 · 7 citationsarXiv
Mesh-TensorFlow: Deep Learning for Supercomputers1811.02084v1 · Noam Shazeer, Youlong Cheng, Niki Parmar et al.2018 · 51 citationsarXiv
Weakly Supervised Grammatical Error Correction using Iterative Decoding1811.01710v1 · Jared Lichtarge, Christopher Alberti, Shankar Kumar et al.2018 · 18 citationsarXiv
Image Transformer1802.05751v3 · Niki Parmar, Ashish Vaswani, Jakob Uszkoreit et al.2018 · 227 citationsarXiv
Fast Decoding in Sequence Models using Discrete Latent Variables1803.03382v6 · \Lukasz Kaiser, Aurko Roy, Ashish Vaswani et al.2018 · 169 citationsarXiv
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost1804.04235v1 · Noam Shazeer, Mitchell Stern2018 · 162 citationsarXiv
Tensor2Tensor for Neural Machine Translation1803.07416v1 · Ashish Vaswani, Samy Bengio, Eugene Brevdo et al.2018 · 14 citationsarXiv
Generating Wikipedia by Summarizing Long Sequences1801.10198v1 · Peter J. Liu, Mohammad Saleh, Etienne Pot et al.2018 · 74 citationsarXiv
One Model To Learn Them All1706.05137v1 · Lukasz Kaiser, Aidan N. Gomez, Noam Shazeer et al.2017 · 255 citationsarXiv
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer1701.06538v1 · Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz et al.2017 · 261 citationsarXiv
NN-grams: Unifying neural network and n-gram language models for Speech Recognition1606.07470v1 · Babak Damavandi, Shankar Kumar, Noam Shazeer et al.2016 · 9 citationsarXiv
Exploring the Limits of Language Modeling1602.02410v2 · Rafal Jozefowicz, Oriol Vinyals, Mike Schuster et al.2016 · 884 citationsarXivon Valency
Swivel: Improving Embeddings by Noticing What's Missing1602.02215v1 · Noam Shazeer, Ryan Doherty, Colin Evans et al.2016 · 56 citationsarXiv
End-to-End Text-Dependent Speaker Verification1509.08062v1 · Georg Heigold, Ignacio Moreno, Samy Bengio et al.2015 · 541 citationsarXiv
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks1506.03099v3 · Samy Bengio, Oriol Vinyals, Navdeep Jaitly et al.2015 · 1,672 citationsarXiv
Skip-gram Language Modeling Using Sparse Non-negative Matrix Probability Estimation1412.1454v2 · Noam Shazeer, Joris Pelemans, Ciprian Chelba2014 · 11 citationsarXiv
Variational Program Inference1006.0991v1 · Georges Harik, Noam Shazeer2010 · 3 citationsarXiv
Career total: 56 works. 36 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.