RG

Roger Grosse

cs.LGstat.MLcs.AIcs.CLmath.OCcs.LOcs.CVcs.GTcs.SEstat.CO

On Valency

published · living versions
W_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version

Preprints & journals

73 papers in the corpus · 2012–2025
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization2509.03378v10 · Wu Lin, Scott C. Lowe, Felix Dangel et al.2025 · 0 citationsarXiv
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference2510.21184v1 · Stephen Zhao, Aidan Li, Rob Brekelmans et al.2025 · 0 citationsarXiv
Better Training Data Attribution via Better Inverse Hessian-Vector Products2507.14740v1 · Andrew Wang, Elisa Nguyen, Runshi Yang et al.2025 · 0 citationsarXiv
Through a Steerable Lens: Magnifying Neural Network Interpretability via Phase-Based Extrapolation2506.02300v3 · Farzaneh Mahdisoltani, Saeed Mahdisoltani, Roger B. Grosse et al.2025 · 0 citationsarXiv
Spectral-factorized Positive-definite Curvature Learning for NN Training2502.06268v3 · Wu Lin, Felix Dangel, Runa Eschenhagen et al.2025 · 0 citationsarXiv
Forecasting Rare Language Model Behaviors2502.16797v1 · Erik Jones, Meg Tong, Jesse Mu et al.2025 · 0 citationsarXiv
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data2406.14546v3 · Johannes Treutlein, Dami Choi, Jan Betley et al.2024 · 5 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Measuring Stochastic Data Complexity with Boltzmann Influence Functions2406.02745v2 · Nathan Ng, Roger Grosse, Marzyeh Ghassemi2024 · 0 citationsarXiv
Training Data Attribution via Approximate Unrolled Differentiation2405.12186v2 · Juhan Bae, Wu Lin, Jonathan Lorraine et al.2024 · 1 citationarXiv
Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo2404.17546v1 · Stephen Zhao, Rob Brekelmans, Alireza Makhzani et al.2024 · 1 citationarXiv
Improving Mutual Information Estimation with Annealed and Energy-Based Bounds2303.06992v1 · Rob Brekelmans, Sicong Huang, Marzyeh Ghassemi et al.2023 · 1 citationICLR 2022 https://openreview.net/forum?id=T0B9AoM_bFg
REFACTOR: Learning to Extract Theorems from Proofs2402.17032v1 · Jin Peng Zhou, Yuhuai Wu, Qiyang Li et al.2024 · 0 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Similarity-based cooperative equilibrium2211.14468v2 · Caspar Oesterheld, Johannes Treutlein, Roger Grosse et al.2022 · 1 citationarXiv
TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) Pipeline for Musical Timbre Transfer1811.09620v3 · Sicong Huang, Qiyang Li, Cem Anil et al.2018 · 75 citationsICLR 2019
Multi-Rate VAE: Train Once, Get the Full Rate-Distortion Curve2212.03905v2 · Juhan Bae, Michael R. Zhang, Michael Ruan et al.2022 · 3 citationsarXiv
Efficient Parametric Approximations of Neural Network Function Space Distance2302.03519v2 · Nikita Dhawan, Sicong Huang, Juhan Bae et al.2023 · 0 citationsarXiv
On Implicit Bias in Overparameterized Bilevel Optimization2212.14032v1 · Paul Vicol, Jonathan Lorraine, Fabian Pedregosa et al.2022 · 3 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Toy Models of Superposition2209.10652v1 · Nelson Elhage, Tristan Hume, Catherine Olsson et al.2022 · 49 citationsarXiv
Learning Branching Heuristics for Propositional Model Counting2007.03204v2 · Pashootan Vaezipoor, Gil Lederman, Yuhuai Wu et al.2020 · 9 citations35(14), 2021, 12427-12435
LIME: Learning Inductive Bias for Primitives of Mathematical Reasoning2101.06223v2 · Yuhuai Wu, Markus Rabe, Wenda Li et al.2021 · 9 citationsarXiv
Amortized Proximal Optimization2203.00089v1 · Juhan Bae, Paul Vicol, Jeff Z. HaoChen et al.2022 · 2 citationsarXiv
Near-optimal Local Convergence of Alternating Gradient Descent-Ascent for Minimax Optimization2102.09468v3 · Guodong Zhang, Yuanhao Wang, Laurent Lessard et al.2021 · 6 citationsarXiv
Understanding and Mitigating Exploding Inverses in Invertible Neural Networks2006.09347v2 · Jens Behrmann, Paul Vicol, Kuan-Chieh Wang et al.2020 · 32 citationsarXiv
Differentiable Annealed Importance Sampling and the Perils of Gradient Noise2107.10211v2 · Guodong Zhang, Kyle Hsu, Jianing Li et al.2021 · 4 citationsarXiv
Regularized linear autoencoders recover the principal components, eventually2007.06731v2 · Xuchan Bao, James Lucas, Sushant Sachdeva et al.2020 · 14 citationsAdvances in Neural Information Processing Systems 33 (NeurIPS 2020)
Learning to Give Checkable Answers with Prover-Verifier Games2108.12099v1 · Cem Anil, Guodong Zhang, Yuhuai Wu et al.2021 · 0 citationsarXiv
Scalable Variational Gaussian Processes via Harmonic Kernel Decomposition2106.05992v1 · Shengyang Sun, Jiaxin Shi, Andrew Gordon Wilson et al.2021 · 1 citationarXiv
A Unified Analysis of First-Order Methods for Smooth Games via Integral Quadratic Constraints2009.11359v4 · Guodong Zhang, Xuchan Bao, Laurent Lessard et al.2020 · 12 citationsarXiv
Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes2104.11044v2 · James Lucas, Juhan Bae, Michael R. Zhang et al.2021 · 7 citationsarXiv
INT: An Inequality Benchmark for Evaluating Generalization in Theorem Proving2007.02924v2 · Yuhuai Wu, Albert Qiaochu Jiang, Jimmy Ba et al.2020 · 20 citationsarXiv
When Does Preconditioning Help or Hurt Generalization?2006.10732v4 · Shun-ichi Amari, Jimmy Ba, Roger Grosse et al.2020 · 11 citationsarXiv
Evaluating Lossy Compression Rates of Deep Generative Models2008.06653v1 · Sicong Huang, Alireza Makhzani, Yanshuai Cao et al.2020 · 11 citationsarXiv
Picking Winning Tickets Before Training by Preserving Gradient Flow2002.07376v2 · Chaoqi Wang, Guodong Zhang, Roger Grosse2020 · 138 citationsIn Proceedings of the 8th International Conference on Learning Representations (ICLR), 2020
Optimizing Neural Networks with Kronecker-factored Approximate Curvature1503.05671v7 · James Martens, Roger Grosse2015 · 490 citationsarXiv
Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks1911.00937v2 · Qiyang Li, Saminul Haque, Cem Anil et al.2019 · 40 citationsarXiv
Don't Blame the ELBO! A Linear VAE Perspective on Posterior Collapse1911.02469v1 · James Lucas, George Tucker, Roger Grosse et al.2019 · 18 citationsarXiv
Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks1905.10961v2 · Guodong Zhang, James Martens, Roger Grosse2019 · 56 citationsarXiv
Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model1907.04164v2 · Guodong Zhang, Lala Li, Zachary Nado et al.2019 · 26 citationsarXiv
Sorting out Lipschitz function approximation1811.05381v2 · Cem Anil, James Lucas, Roger Grosse2018 · 70 citationsarXiv
EigenDamage: Structured Pruning in the Kronecker-Factored Eigenbasis1905.05934v1 · Chaoqi Wang, Roger Grosse, Sanja Fidler et al.2019 · 30 citationsarXiv
Aggregated Momentum: Stability Through Passive Damping1804.00325v3 · James Lucas, Shengyang Sun, Richard Zemel et al.2018 · 7 citationsInternational Conference on Learning Representations, 2019
Isolating Sources of Disentanglement in Variational Autoencoders1802.04942v5 · Ricky T. Q. Chen, Xuechen Li, Roger Grosse et al.2018 · 181 citationsarXiv
Functional Variational Bayesian Neural Networks1903.05779v1 · Shengyang Sun, Guodong Zhang, Jiaxin Shi et al.2019 · 137 citationsarXiv
Self-Tuning Networks: Bilevel Optimization of Hyperparameters using Structured Best-Response Functions1903.03088v1 · Matthew MacKay, Paul Vicol, Jon Lorraine et al.2019 · 81 citationsarXiv
Career total: 102 works. 73 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.