YC
Yuan Cao
cs.LGcs.CLstat.MLcs.AImath.OCcs.SDeess.AScs.CVmath.STstat.TH
On Valency
published · living versionsW_zb7r6dq8·v1 · currentpublished
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
with Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le +26
1 version
Preprints & journals
47 papers in the corpus · 2016–2025Structure-Aware Automatic Channel Pruning by Searching with Graph Embedding2506.11469v1 · Zifan Liu, Yuan Cao, Yanwei Yu et al.2025 · 0 citationsarXiv
Gemini: A Family of Highly Capable Multimodal Models2312.11805v5 · Gemini Team Google: Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et al.2023 · 838 citationsarXiv
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks2504.08628v1 · Chenyang Zhang, Peifeng Gao, Difan Zou et al.2025 · 0 citationsarXiv
Transformer Learns Optimal Variable Selection in Group-Sparse Classification2504.08638v1 · Chenyang Zhang, Xuran Meng, Yuan Cao2025 · 0 citationsarXiv
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context2411.10830v1 · Zihao Li, Yuan Cao, Cheng Gao et al.2024 · 0 citationsarXiv
Global Convergence in Training Large-Scale Transformers2410.23610v1 · Cheng Gao, Yuan Cao, Zihao Li et al.2024 · 1 citationarXiv
Improving Fast Adversarial Training via Self-Knowledge Guidance2409.17589v1 · Chengze Jiang, Junkai Wang, Minjing Dong et al.2024 · 11 citationsarXiv
Retrieval Augmented End-to-End Spoken Dialog Models2402.01828v1 · Mingqiu Wang, Izhak Shafran, Hagen Soltau et al.2024 · 14 citationsProc. ICASSP 2024
SLM: Bridge the thin gap between speech and text foundation models2310.00230v1 · Mingqiu Wang, Wei Han, Izhak Shafran et al.2023 · 30 citationsarXiv
Learn to Cluster Faces with Better Subgraphs2304.10831v1 · Yuan Cao, Di Jiang, Guanqun Hou et al.2023 · 2 citationsarXiv
Show, Don't Tell: Demonstrations Outperform Descriptions for Schema-Guided Task-Oriented Dialogue2204.04327v2 · Raghav Gupta, Harrison Lee, Jeffrey Zhao et al.2022 · 19 citationsIn Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4541-4549, Seattle, United States. Association for Computational Linguistics
SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue Systems2110.06800v3 · Harrison Lee, Raghav Gupta, Abhinav Rastogi et al.2021 · 17 citationsLee, H., Gupta, R., Rastogi, A., Cao, Y., Zhang, B., & Wu, Y. (2022). SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue Systems. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10), 10938-10946
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision2108.10904v3 · Zirui Wang, Jiahui Yu, Adams Wei Yu et al.2021 · 343 citationsarXiv
Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation2203.07627v1 · Yong Cheng, Ankur Bapna, Orhan Firat et al.2022 · 13 citationsarXiv
Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures2104.13628v2 · Yuan Cao, Quanquan Gu, Mikhail Belkin2021 · 12 citationsarXiv
Efficient and Private Federated Learning with Partially Trainable Networks2110.03450v2 · Hakim Sidahmed, Zheng Xu, Ankush Garg et al.2021 · 11 citationsarXiv
Towards Zero-Label Language Learning2109.09193v1 · Zirui Wang, Adams Wei Yu, Orhan Firat et al.2021 · 46 citationsarXiv
Effective Sequence-to-Sequence Dialogue State Tracking2108.13990v2 · Jeffrey Zhao, Mahdis Mahdieh, Ye Zhang et al.2021 · 28 citationsarXiv
Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization2108.11371v1 · Difan Zou, Yuan Cao, Yuanzhi Li et al.2021 · 6 citationsarXiv
Your GAN is Secretly an Energy-based Model and You Should use Discriminator Driven Latent Sampling2003.06060v3 · Tong Che, Ruixiang Zhang, Jascha Sohl-Dickstein et al.2020 · 52 citationsarXiv
Improving Longer-range Dialogue State Tracking2103.00109v2 · Ye Zhang, Yuan Cao, Mahdis Mahdieh et al.2021 · 3 citationsarXiv
Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise2101.01152v3 · Spencer Frei, Yuan Cao, Quanquan Gu2021 · 2 citationsarXiv
Towards NNGP-guided Neural Architecture Search2011.06006v1 · Daniel S. Park, Jaehoon Lee, Daiyi Peng et al.2020 · 13 citationsarXiv
Rapid Domain Adaptation for Machine Translation with Monolingual Data2010.12652v1 · Mahdis Mahdieh, Mia Xu Chen, Yuan Cao et al.2020 · 2 citationsarXiv
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior2010.11054v1 · Jiaming Luo, Frederik Hartmann, Enrico Santus et al.2020 · 20 citationsarXiv
Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models2010.05874v1 · Zirui Wang, Yulia Tsvetkov, Orhan Firat et al.2020 · 75 citationsarXiv
A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks2002.04026v2 · Zixiang Chen, Yuan Cao, Quanquan Gu et al.2020 · 29 citationsarXiv
Towards Understanding the Spectral Bias of Deep Learning1912.01198v3 · Yuan Cao, Zhiying Fang, Yue Wu et al.2019 · 153 citationsarXiv
Agnostic Learning of a Single Neuron with Gradient Descent2005.14426v3 · Spencer Frei, Yuan Cao, Quanquan Gu2020 · 16 citationsarXiv
Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks1806.06763v3 · Jinghui Chen, Dongruo Zhou, Yiqi Tang et al.2018 · 142 citationsarXiv
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation2005.04816v1 · Aditya Siddhant, Ankur Bapna, Yuan Cao et al.2020 · 78 citationsACL 2020
Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis2002.03785v1 · Guangzhi Sun, Yu Zhang, Ron J. Weiss et al.2020 · 100 citationsarXiv
Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior2002.03788v1 · Guangzhi Sun, Yu Zhang, Ron J. Weiss et al.2020 · 92 citationsarXiv
Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks1902.01384v4 · Yuan Cao, Quanquan Gu2019 · 119 citationsarXiv
Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks1905.13210v3 · Yuan Cao, Quanquan Gu2019 · 184 citationsarXiv
Tight Sample Complexity of Learning One-hidden-layer Convolutional Neural Networks1911.05059v1 · Yuan Cao, Quanquan Gu2019 · 15 citationsarXiv
Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks1910.02934v1 · Spencer Frei, Yuan Cao, Quanquan Gu2019 · 18 citationsarXiv
Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges1907.05019v1 · Naveen Arivazhagan, Ankur Bapna, Orhan Firat et al.2019 · 287 citationsarXiv
Neural Decipherment via Minimum-Cost Flow: from Ugaritic to Linear B1906.06718v1 · Jiaming Luo, Yuan Cao, Regina Barzilay2019 · 32 citationsarXiv
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling1902.08295v1 · Jonathan Shen, Patrick Nguyen, Yonghui Wu et al.2019 · 181 citationsarXiv
Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation1811.02050v2 · Ye Jia, Melvin Johnson, Wolfgang Macherey et al.2018 · 158 citationsarXiv
Hierarchical Generative Modeling for Controllable Speech Synthesis1810.07217v2 · Wei-Ning Hsu, Yu Zhang, Ron J. Weiss et al.2018 · 43 citationsarXiv
Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks1811.08888v3 · Difan Zou, Yuan Cao, Dongruo Zhou et al.2018 · 186 citationsarXiv
Training Deeper Neural Machine Translation Models with Transparent Attention1808.07561v2 · Ankur Bapna, Mia Xu Chen, Orhan Firat et al.2018 · 158 citationsarXiv
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation1609.08144v2 · Yonghui Wu, Mike Schuster, Zhifeng Chen et al.2016 · 5,573 citationsarXivon Valency
Patient performance-based plan parameter optimization for prostate cancer in tomotherapy.25869936 · Cao, Yuan Jie, Lee, Suk, Chang, Kyung Hwan et al.2016 · 14 citationsMedical dosimetry : official journal of the American Association of Medical Dosimetrists. 2015;40(4):285-9
Development of a 3D optical scanner for evaluating patient-specific dose distributions.26048682 · Chang, Kyung Hwan, Lee, Suk, Jung, Hong et al.2016 · 10 citationsPhysica medica : PM : an international journal devoted to the applications of physics to medicine and biology : official journal of the Italian Association of Biomedical Physics (AIFB). 2015;31(5):553-9
Career total: 97 works. 47 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.