AV
Ashish Vaswani
cs.LGcs.CLstat.MLcs.AIcs.CVeess.ASbioinformaticsCognitive NeuroscienceCognitive Psychologycs.DC
On Valency
published · living versionsW_qeugbgvj·v1 · currentpublished
Preprints & journals
28 papers in the corpus · 2016–2026A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio2605.28975v1 · Ali Shehper, Ashish Vaswani2026 · 0 citationsarXiv
Essential-Web v1.0: 24T tokens of organized web data2506.14111v2 · Essential AI: Andrew Hojel, Michael Pust, Tim Romanski et al.2025 · 0 citationsarXiv
Practical Efficiency of Muon for Pretraining2505.02222v4 · Essential AI: Ishaan Shah, Anthony M. Polloreno, Karl Stratos et al.2025 · 0 citationsarXiv
Rethinking Reflection in Pre-Training2504.04022v1 · Essential AI: Darsh J Shah, Peter Rushton, Somanshu Singla et al.2025 · 0 citationsarXiv
Attention Is All You Need1706.03762v7 · Ashish Vaswani, Noam Shazeer, Niki Parmar et al.2017 · 26,892 citationsarXivon Valency
DeepConsensus improves the accuracy of sequences with a gap-aware sequence transformer.36050551 · Baid, Gunjan, Cook, Daniel E, Shafin, Kishwar et al.2023 · 161 citationsNature biotechnology. 2023;41(2):232-238
The Efficiency Misnomer2110.12894v2 · Mostafa Dehghani, Anurag Arnab, Lucas Beyer et al.2021 · 0 citationsarXiv
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers2109.10686v2 · Yi Tay, Mostafa Dehghani, Jinfeng Rao et al.2021 · 56 citationsarXiv
DeepConsensus: Gap-Aware Sequence Transformers for Sequence Correction10.1101/2021.08.31.458403v1 · Baid, G., Cook, D. E., Shafin, K. et al.2021 · 10 citationsbioRxiv
Bottleneck Transformers for Visual Recognition2101.11605v2 · Aravind Srinivas, Tsung-Yi Lin, Niki Parmar et al.2021 · 1,304 citationsarXiv
Scaling Local Self-Attention for Parameter Efficient Visual Backbones2103.12731v3 · Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas et al.2021 · 461 citationsarXiv
Simple and Efficient ways to Improve REALM2104.08710v1 · Vidhisha Balachandran, Ashish Vaswani, Yulia Tsvetkov et al.2021 · 4 citationsarXiv
Efficient Content-Based Sparse Attention with Routing Transformers2003.05997v5 · Aurko Roy, Mohammad Saffar, Ashish Vaswani et al.2020 · 517 citationsarXiv
Attention Augmented Convolutional Networks1904.09925v5 · Irwan Bello, Barret Zoph, Ashish Vaswani et al.2019 · 1,249 citationsarXiv
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation1905.12255v3 · Vihan Jain, Gabriel Magalhaes, Alexander Ku et al.2019 · 169 citationsarXiv
Stand-Alone Self-Attention in Vision Models1906.05909v1 · Prajit Ramachandran, Niki Parmar, Ashish Vaswani et al.2019 · 216 citationsarXiv
Music Transformer1809.04281v3 · Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit et al.2018 · 73 citationsarXiv
Mesh-TensorFlow: Deep Learning for Supercomputers1811.02084v1 · Noam Shazeer, Youlong Cheng, Niki Parmar et al.2018 · 51 citationsarXiv
Relational inductive biases, deep learning, and graph networks1806.01261v3 · Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst et al.2018 · 2,333 citationsarXiv
Theory and Experiments on Vector Quantized Autoencoders1805.11063v2 · Aurko Roy, Ashish Vaswani, Arvind Neelakantan et al.2018 · 54 citationsarXiv
Image Transformer1802.05751v3 · Niki Parmar, Ashish Vaswani, Jakob Uszkoreit et al.2018 · 227 citationsarXiv
Decoding the neural representation of story meanings across languages.28940969 · Dehghani, Morteza, Boghrati, Reihane, Man, Kingson et al.2018 · 94 citationsHuman brain mapping. 2017;38(12):6096-6106
Fast Decoding in Sequence Models using Discrete Latent Variables1803.03382v6 · \Lukasz Kaiser, Aurko Roy, Ashish Vaswani et al.2018 · 169 citationsarXiv
Self-Attention with Relative Position Representations1803.02155v2 · Peter Shaw, Jakob Uszkoreit, Ashish Vaswani2018 · 2,526 citationsarXiv
Tensor2Tensor for Neural Machine Translation1803.07416v1 · Ashish Vaswani, Samy Bengio, Eugene Brevdo et al.2018 · 14 citationsarXiv
One Model To Learn Them All1706.05137v1 · Lukasz Kaiser, Aidan N. Gomez, Noam Shazeer et al.2017 · 255 citationsarXiv
Decoding the Neural Representation of Story Meanings across Languagesqrpp3_v1 · Morteza Dehghani, Reihane Boghrati, Kingson Man et al.2017 · 85 citationsPsyArXiv
Unsupervised Neural Hidden Markov Models1609.09007v1 · Ke Tran, Yonatan Bisk, Ashish Vaswani et al.2016 · 65 citationsarXiv
Career total: 61 works. 28 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.