SB
Samy Bengio
cs.LGstat.MLcs.AIcs.CLcs.CVArtificial IntelligenceAlgorithmscs.SDcs.CRcs.NE
On Valency
published · living versionsW_mnxekqpq·v1 · currentpublished
Zero-Shot Learning by Convex Combination of Semantic Embeddings
with Mohammad Norouzi, Tomas Mikolov, Yoram Singer, Jonathon Shlens +3
1 version
Preprints & journals
83 papers in the corpus · 2002–2026The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models2609.21509v2 · Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella et al.2026 · 0 citationsarXiv
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context2603.01357v1 · Zidi Xiu, David Q. Sun, Kevin Cheng et al.2026 · 0 citationsarXiv
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity2506.06941v3 · Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh et al.2025 · 29 citationsarXiv
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models2410.05229v2 · Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi et al.2024 · 55 citationsarXiv
Boolformer: Symbolic Regression of Logic Functions with Transformers2309.12207v2 · St'ephane d'Ascoli, Arthur Renard, Vassilis Papadopoulos et al.2023 · 1 citationarXiv
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?2502.18435v3 · Yizhe Zhang, Richard Bai, Zijin Gu et al.2025 · 0 citationsarXiv
Chain-of-Sketch: Enabling Global Visual Reasoning2410.08165v2 · Aryo Lotfi, Enrico Fini, Samy Bengio et al.2024 · 0 citationsarXiv
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining2504.02107v3 · Jeffrey Li, Mohammadreza Armandpour, Iman Mirzadeh et al.2025 · 2 citationsarXiv
Generalization on the Unseen, Logic Reasoning and Degree Curriculum2301.13105v3 · Emmanuel Abbe, Samy Bengio, Aryo Lotfi et al.2023 · 3 citationsarXiv
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad2406.06467v3 · Emmanuel Abbe, Samy Bengio, Aryo Lotfi et al.2024 · 1 citationarXiv
When can transformers reason with abstract symbols?2310.09753v2 · Enric Boix-Adsera, Omid Saremi, Emmanuel Abbe et al.2023 · 2 citationsarXiv
Transformers learn through gradual rank increase2306.07042v2 · Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe et al.2023 · 6 citationsarXiv
What Algorithms can Transformers Learn? A Study in Length Generalization2310.16028v1 · Hattie Zhou, Arwen Bradley, Etai Littwin et al.2023 · 8 citationsarXiv
Adaptivity and Modularity for Efficient Generalization Over Task Complexity2310.08866v1 · Samira Abnar, Omid Saremi, Laurent Dinh et al.2023 · 0 citationsarXiv
Continuous Pseudo-Labeling from the Start2210.08711v2 · Dan Berrebbi, Ronan Collobert, Samy Bengio et al.2022 · 6 citationsarXiv
Continuous Soft Pseudo-Labeling in ASR2211.06007v2 · Tatiana Likhomanenko, Ronan Collobert, Navdeep Jaitly et al.2022 · 3 citationsarXiv
Learning to Reason with Neural Networks: Generalization, Unseen Data and Boolean Measures2205.13647v2 · Emmanuel Abbe, Samy Bengio, Elisabetta Cornacchia et al.2022 · 3 citationsarXiv
Are All Layers Created Equal?1902.01996v4 · Chiyuan Zhang, Samy Bengio, Yoram Singer2019 · 106 citationsarXiv
Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization2107.12580v2 · Chiyuan Zhang, Maithra Raghu, Jon Kleinberg et al.2021 · 2 citationsarXiv
Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding2106.02795v3 · Yang Li, Si Si, Gang Li et al.2021 · 31 citationsarXiv
Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference Loss2102.09808v4 · Michael L. Iuzzolino, Michael C. Mozer, Samy Bengio2021 · 5 citationsarXiv
Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards1907.10247v3 · Yijie Guo, Jongwook Choi, Marcin Moczulski et al.2019 · 14 citationsarXiv
Characterising Bias in Compressed Models2010.03058v2 · Sara Hooker, Nyalleng Moorosi, Gregory Clark et al.2020 · 44 citationsarXiv
NeurIPS 2020 Competition: Predicting Generalization in Deep Learning2012.07976v1 · Yiding Jiang, Pierre Foret, Scott Yak et al.2020 · 21 citationsarXiv
Data Augmentation via Structured Adversarial Perturbations2011.03010v1 · Calvin Luo, Hossein Mobahi, Samy Bengio2020 · 1 citationarXiv
Neural Networks Trained on Natural Scenes Exhibit Gestalt Closure1903.01069v4 · Been Kim, Emily Reif, Martin Wattenberg et al.2019 · 43 citationsarXiv
Area Attention1810.10126v7 · Yang Li, Lukasz Kaiser, Samy Bengio et al.2018 · 0 citationsICML 2019
Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML1909.09157v2 · Aniruddh Raghu, Maithra Raghu, Samy Bengio et al.2019 · 94 citationsarXiv
Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders2001.05308v1 · Yang Li, Julien Amelot, Xin Zhou et al.2020 · 11 citationsarXiv
Identity Crisis: Memorization and Generalization under Extreme Overparameterization1902.04698v4 · Chiyuan Zhang, Samy Bengio, Moritz Hardt et al.2019 · 50 citationsarXiv
Fantastic Generalization Measures and Where to Find Them1912.02178v1 · Yiding Jiang, Behnam Neyshabur, Hossein Mobahi et al.2019 · 225 citationsarXiv
Transfusion: Understanding Transfer Learning for Medical Imaging1902.07208v3 · Maithra Raghu, Chiyuan Zhang, Jon Kleinberg et al.2019 · 756 citationsarXivon Valency
Parallel Scheduled Sampling1906.04331v2 · Daniel Duckworth, Arvind Neelakantan, Ben Goodrich et al.2019 · 12 citationsarXiv
Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks1905.07953v2 · Wei-Lin Chiang, Xuanqing Liu, Si Si et al.2019 · 1,399 citationsarXiv
Predicting the Generalization Gap in Deep Networks with Margin Distributions1810.00113v2 · Yiding Jiang, Dilip Krishnan, Hossein Mobahi et al.2018 · 123 citationsarXiv
A Closed-Form Learned Pooling for Deep Classification Networks1906.03808v1 · Vighnesh Birodkar, Hossein Mobahi, Dilip Krishnan et al.2019 · 0 citationsarXiv
You Look Twice: GaterNet for Dynamic Filter Selection in CNNs1811.11205v2 · Zhourong Chen, Yang Li, Samy Bengio et al.2018 · 107 citationsarXiv
Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need1901.11409v1 · Vighnesh Birodkar, Hossein Mobahi, Samy Bengio2019 · 33 citationsarXiv
Large Margin Deep Networks for Classification1803.05598v2 · Gamaleldin F. Elsayed, Dilip Krishnan, Hossein Mobahi et al.2018 · 14 citationsarXiv
Content preserving text generation with attribute controls1811.01135v1 · Lajanugen Logeswaran, Honglak Lee, Samy Bengio2018 · 81 citationsarXiv
Insights on representational similarity in neural networks with canonical correlation1806.05759v3 · Ari S. Morcos, Maithra Raghu, Samy Bengio2018 · 148 citationsarXiv
Time-Dependent Representation for Neural Event Sequence Prediction1708.00065v4 · Yang Li, Nan Du, Samy Bengio2017 · 5 citationsarXiv
Fast Decoding in Sequence Models using Discrete Latent Variables1803.03382v6 · \Lukasz Kaiser, Aurko Roy, Ashish Vaswani et al.2018 · 169 citationsarXiv
A Study on Overfitting in Deep Reinforcement Learning1804.06893v2 · Chiyuan Zhang, Oriol Vinyals, Remi Munos et al.2018 · 231 citationsarXiv
Adversarial Attacks and Defences Competition1804.00097v1 · Alexey Kurakin, Ian Goodfellow, Samy Bengio et al.2018 · 329 citationsarXiv
Tensor2Tensor for Neural Machine Translation1803.07416v1 · Ashish Vaswani, Samy Bengio, Eugene Brevdo et al.2018 · 14 citationsarXiv
Predicting Human Performance in Vertical Menu Selection Using Deep Learning1803.05073v1 · Yang Li, Samy Bengio, Gilles Bailly2018 · 33 citationsarXiv
On Using Backpropagation for Speech Texture Generation and Voice Conversion1712.08363v2 · Jan Chorowski, Ron J. Weiss, Rif A. Saurous et al.2017 · 21 citationsarXiv
Discrete Autoencoders for Sequence Models1801.09797v1 · \Lukasz Kaiser, Samy Bengio2018 · 34 citationsarXiv
Sharp Minima Can Generalize For Deep Nets1703.04933v2 · Laurent Dinh, Razvan Pascanu, Samy Bengio et al.2017 · 134 citationsarXiv
Career total: 389 works. 83 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.