DG
Deep Ganguli
cs.CLcs.AIcs.LGcs.CYcs.CRcs.HCq-bio.NCArtificial IntelligenceBayes Theoremcs.GL
On Valency
published · living versionsW_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version
Preprints & journals
27 papers in the corpus · 2012–2025The impact of advanced AI systems on democracy.41034566 · Summerfield, Christopher, Argyle, Lisa P, Bakker, Michiel et al.2025 · 13 citationsNature human behaviour. 2025;9(12):2420-2430
Implicit encoding of prior probabilities in optimal neural populations.25356064 · Ganguli, Deep, Simoncelli, Eero P2025 · 93 citationsAdvances in neural information processing systems. 2010;2010:658-666
Toward an Evaluation Science for Generative AI Systems2503.05336v3 · Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach et al.2025 · 3 citationsarXiv
Clio: Privacy-Preserving Insights into Real-World AI Use2412.13678v1 · Alex Tamkin, Miles McCain, Kunal Handa et al.2024 · 8 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
How will advanced AI systems impact democracy?2409.06729v1 · Christopher Summerfield, Lisa Argyle, Michiel Bakker et al.2024 · 5 citationsarXiv
Collective Constitutional AI: Aligning a Language Model with Public Input2406.07814v1 · Saffron Huang, Divya Siddarth, Liane Lovitt et al.2024 · 75 citationsProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1395-1417
Towards Measuring the Representation of Subjective Global Opinions in Language Models2306.16388v2 · Esin Durmus, Karina Nguyen, Thomas I. Liao et al.2023 · 43 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Evaluating and Mitigating Discrimination in Language Model Decisions2312.03689v1 · Alex Tamkin, Amanda Askell, Liane Lovitt et al.2023 · 10 citationsarXiv
Report of the 1st Workshop on Generative AI and Law2311.06477v3 · A. Feder Cooper, Katherine Lee, James Grimmelmann et al.2023 · 6 citationsarXiv
Opportunities and Risks of LLMs for Scalable Deliberation with Polis2306.11932v1 · Christopher T. Small, Ivan Vendrov, Esin Durmus et al.2023 · 20 citationsarXiv
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models2206.04615v3 · Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al.2022 · 565 citationsTransactions on Machine Learning Research, May/2022, https://openreview.net/forum?id=uyTL5Bvosj
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
In-context Learning and Induction Heads2209.11895v1 · Catherine Olsson, Nelson Elhage, Neel Nanda et al.2022 · 87 citationsarXivon Valency
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
A General Language Assistant as a Laboratory for Alignment2112.00861v3 · Amanda Askell, Yuntao Bai, Anna Chen et al.2021 · 27 citationsarXiv
The AI Index 2021 Annual Report2103.06312v1 · Daniel Zhang, Saurabh Mishra, Erik Brynjolfsson et al.2021 · 9 citationsarXiv
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models2102.02503v1 · Alex Tamkin, Miles Brundage, Jack Clark et al.2021 · 129 citationsarXiv
Neural and perceptual signatures of efficient sensory coding1603.00058v1 · Deep Ganguli, Eero P. Simoncelli2016 · 21 citationsarXiv
Efficient sensory encoding and Bayesian inference with heterogeneous neural populations.25058702 · Ganguli, Deep, Simoncelli, Eero P2015 · 249 citationsNeural computation. 2014;26(10):2103-34
Implicit embedding of prior probabilities in optimally efficient neural populations1209.5006v1 · Deep Ganguli, Eero Simoncelli2012 · 1 citationarXiv
Career total: 32 works. 27 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.