AA
Amanda Askell
cs.CLcs.AIcs.LGcs.CYcs.CRstat.MLcs.CVcs.HCcs.SE
On Valency
published · living versionsW_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version
Preprints & journals
23 papers in the corpus · 2019–2025Towards Understanding Sycophancy in Language Models2310.13548v4 · Mrinank Sharma, Meg Tong, Tomasz Korbak et al.2023 · 132 citationsarXiv
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming2501.18837v1 · Mrinank Sharma, Meg Tong, Jesse Mu et al.2025 · 7 citationsarXiv
Towards Measuring the Representation of Subjective Global Opinions in Language Models2306.16388v2 · Esin Durmus, Karina Nguyen, Thomas I. Liao et al.2023 · 43 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Evaluating and Mitigating Discrimination in Language Model Decisions2312.03689v1 · Alex Tamkin, Amanda Askell, Liane Lovitt et al.2023 · 10 citationsarXiv
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models2206.04615v3 · Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al.2022 · 565 citationsTransactions on Machine Learning Research, May/2022, https://openreview.net/forum?id=uyTL5Bvosj
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
In-context Learning and Induction Heads2209.11895v1 · Catherine Olsson, Nelson Elhage, Neel Nanda et al.2022 · 87 citationsarXivon Valency
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
Training language models to follow instructions with human feedback2203.02155v1 · Long Ouyang, Jeff Wu, Xu Jiang et al.2022 · 5,565 citationsarXiv
A General Language Assistant as a Laboratory for Alignment2112.00861v3 · Amanda Askell, Yuntao Bai, Anna Chen et al.2021 · 27 citationsarXiv
Learning Transferable Visual Models From Natural Language Supervision2103.00020v1 · Alec Radford, Jong Wook Kim, Chris Hallacy et al.2021 · 5,285 citationsarXiv
Language Models are Few-Shot Learners2005.14165v4 · Tom B. Brown, Benjamin Mann, Nick Ryder et al.2020 · 2,966 citationsarXiv
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims2004.07213v2 · Miles Brundage, Shahar Avin, Jasmine Wang et al.2020 · 301 citationsarXiv
Release Strategies and the Social Impacts of Language Models1908.09203v2 · Irene Solaiman, Miles Brundage, Jack Clark et al.2019 · 282 citationsarXiv
The Role of Cooperation in Responsible AI Development1907.04534v1 · Amanda Askell, Miles Brundage, Gillian Hadfield2019 · 46 citationsarXiv
Career total: 32 works. 23 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.