DD

David Duvenaud

cs.LGstat.MLcs.AIcs.CLcs.CVcs.CRcs.CYcs.NAmath.NAcs.SE

On Valency

published · living versions
W_uutna755·v1 · currentpublished
Alignment faking in large language models
with Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger +15
1 version

Preprints & journals

71 papers in the corpus · 2011–2026
Towards Understanding Linear Word Analogies1810.04882v9 · Kawin Ethayarajh, David Duvenaud, Graeme Hirst2018 · 84 citationsarXiv
Understanding Undesirable Word Embedding Associations1908.06361v2 · Kawin Ethayarajh, David Duvenaud, Graeme Hirst2019 · 127 citationsarXiv
Optimally-Weighted Herding is Bayesian Quadrature1204.1664v3 · Ferenc Husz'ar, David Duvenaud2012 · 26 citationsarXiv
The Artificial Self: Characterising the landscape of AI identity2603.11353v1 · Raymond Douglas, Jan Kulveit, Ondrej Havlicek et al.2026 · 1 citationarXiv
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage2601.19062v1 · Mrinank Sharma, Miles McCain, Raymond Douglas et al.2026 · 2 citationsarXiv
A Definition of AGI2510.18212v3 · Dan Hendrycks, Dawn Song, Christian Szegedy et al.2025 · 3 citationsarXiv
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value2512.03399v1 · Joe Edelman, Tan Zhi-Xuan, Ryan Lowe et al.2025 · 0 citationsarXiv
Towards Understanding Sycophancy in Language Models2310.13548v4 · Mrinank Sharma, Meg Tong, Tomasz Korbak et al.2023 · 132 citationsarXiv
JoLT: Joint Probabilistic Predictions on Tabular Data Using LLMs2502.11877v1 · Aliaksandra Shysheya, John Bronskill, James Requeima et al.2025 · 1 citationarXiv
Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development2501.16946v2 · Jan Kulveit, Raymond Douglas, Nora Ammann et al.2025 · 12 citationsarXiv
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language2405.12856v5 · James Requeima, John Bronskill, Dami Choi et al.2024 · 10 citations38th Conference on Neural Information Processing Systems (NeurIPS 2024)
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs2402.08733v2 · Daniel D. Johnson, Daniel Tarlow, David Duvenaud et al.2024 · 1 citationarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Sorting Out Quantum Monte Carlo2311.05598v1 · Jack Richter-Powell, Luca Thiede, Al'an Asparu-Guzik et al.2023 · 1 citationarXiv
Tools for Verifying Neural Models' Training Data2307.00682v1 · Dami Choi, Yonadav Shavit, David Duvenaud2023 · 5 citationsarXiv
On Implicit Bias in Overparameterized Bilevel Optimization2212.14032v1 · Paul Vicol, Jonathan Lorraine, Fabian Pedregosa et al.2022 · 3 citationsarXiv
Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations2102.06559v4 · Winnie Xu, Ricky T.Q. Chen, Xuechen Li et al.2021 · 15 citationsarXiv
Meta-Learning to Improve Pre-Training2111.01754v1 · Aniruddh Raghu, Jonathan Lorraine, Simon Kornblith et al.2021 · 0 citationsarXiv
No MCMC for me: Amortized sampling for fast and stable training of energy-based models2010.04230v3 · Will Grathwohl, Jacob Kelly, Milad Hashemi et al.2020 · 11 citationsarXiv
Oops I Took A Gradient: Scalable Sampling for Discrete Distributions2102.04509v2 · Will Grathwohl, Kevin Swersky, Milad Hashemi et al.2021 · 8 citationsarXiv
Complex Momentum for Optimization in Games2102.08431v2 · Jonathan Lorraine, David Acuna, Paul Vicol et al.2021 · 0 citationsarXiv
Getting to the Point. Index Sets and Parallelism-Preserving Autodiff for Pointful Array Programming2104.05372v1 · Adam Paszke, Daniel Johnson, David Duvenaud et al.2021 · 40 citationsarXiv
Teaching with Commentaries2011.03037v2 · Aniruddh Raghu, Maithra Raghu, Simon Kornblith et al.2020 · 4 citationsarXiv
Self-Tuning Stochastic Optimization with Curvature-Aware Gradient Filtering2011.04803v1 · Ricky T. Q. Chen, Dami Choi, Lukas Balles et al.2020 · 2 citationsarXiv
What went wrong and when? Instance-wise Feature Importance for Time-series Models2003.02821v3 · Sana Tonekaboni, Shalmali Joshi, Kieran Campbell et al.2020 · 9 citationsarXiv
Learning Differential Equations that are Easy to Solve2007.04504v2 · Jacob Kelly, Jesse Bettencourt, Matthew James Johnson et al.2020 · 55 citationsarXiv
Scalable Gradients for Stochastic Differential Equations2001.01328v6 · Xuechen Li, Ting-Kam Leonard Wong, Ricky T. Q. Chen et al.2020 · 141 citationsarXiv
Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One1912.03263v3 · Will Grathwohl, Kuan-Chieh Wang, Jorn-Henrik Jacobsen et al.2019 · 184 citationsarXiv
Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without Sampling2002.05616v4 · Will Grathwohl, Kuan-Chieh Wang, Jorn-Henrik Jacobsen et al.2020 · 34 citationsarXiv
Residual Flows for Invertible Generative Modeling1906.02735v6 · Ricky T. Q. Chen, Jens Behrmann, David Duvenaud et al.2019 · 196 citationsarXiv
Efficient Graph Generation with Graph Recurrent Attention Networks1910.00760v3 · Renjie Liao, Yujia Li, Yang Song et al.2019 · 168 citationsarXiv
SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable Models2004.00353v2 · Yucen Luo, Alex Beatson, Mohammad Norouzi et al.2020 · 16 citationsarXiv
A Study of Gradient Variance in Deep Learning2007.04532v1 · Fartash Faghri, David Duvenaud, David J. Fleet et al.2020 · 12 citationsarXiv
Neural Ordinary Differential Equations1806.07366v5 · Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt et al.2018 · 558 citationsarXiv
Neural Networks with Cheap Differential Operators1912.03579v1 · Ricky T. Q. Chen, David Duvenaud2019 · 12 citationsarXiv
Optimizing Millions of Hyperparameters by Implicit Differentiation1911.02590v1 · Jonathan Lorraine, Paul Vicol, David Duvenaud2019 · 41 citationsarXiv
Latent ODEs for Irregularly-Sampled Time Series1907.03907v1 · Yulia Rubanova, Ricky T. Q. Chen, David Duvenaud2019 · 145 citationsarXiv
Invertible Residual Networks1811.00995v3 · Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen et al.2018 · 159 citationsProceedings of the International Conference on Machine Learning (ICML), 2019
Isolating Sources of Disentanglement in Variational Autoencoders1802.04942v5 · Ricky T. Q. Chen, Xuechen Li, Roger Grosse et al.2018 · 181 citationsarXiv
Self-Tuning Networks: Bilevel Optimization of Hyperparameters using Structured Best-Response Functions1903.03088v1 · Matthew MacKay, Paul Vicol, Jon Lorraine et al.2019 · 81 citationsarXiv
Explaining Image Classifiers by Counterfactual Generation1807.08024v3 · Chun-Hao Chang, Elliot Creager, Anna Goldenberg et al.2018 · 84 citationsarXiv
FFJORD: Free-form Continuous Dynamics for Scalable Reversible Generative Models1810.01367v3 · Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt et al.2018 · 439 citationsarXiv
Stochastic Combinatorial Ensembles for Defending Against Adversarial Examples1808.06645v2 · George A. Adam, Petr Smirnov, David Duvenaud et al.2018 · 4 citationsarXiv
Design of efficient molecular organic light-emitting diodes by a high-throughput virtual screening and experimental approach.27500805 · Gómez-Bombarelli, Rafael, Aguilera-Iparraguirre, Jorge, Hirzel, Timothy D et al.2018 · 1,147 citationsNature materials. 2016;15(10):1120-7
Scalable Recommender Systems through Recursive Evidence Chains1807.02150v1 · Elias Tragas, Calvin Luo, Maxime Gazeau et al.2018 · 0 citationsarXiv
Compositional inductive biases in function learning.29154187 · Schulz, Eric, Tenenbaum, Joshua B, Duvenaud, David et al.2018 · 110 citationsCognitive psychology. 2017;99:44-79
cDeepbind: A context sensitive deep learning model of RNA-protein binding10.1101/345140v1 · Gandhi, S., Lee, L. J., Delong, A. et al.2018 · 19 citationsbioRxiv
Career total: 105 works. 71 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.