EP
Ethan Perez
cs.CLcs.AIcs.LGstat.MLcs.CRcs.CVcs.CYcs.SEcs.HCcs.IR
On Valency
published · living versionsW_vv4hg27n·v1 · currentpublished
Agentic Misalignment: How LLMs Could Be Insider Threats
with Aengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie +3
1 version
Preprints & journals
60 papers in the corpus · 2016–2026The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?2601.23045v2 · Alexander Hagele, Aryo Pradipta Gema, Henry Sleight et al.2026 · 1 citationarXiv
Unsupervised Elicitation of Language Models2506.10139v2 · Jiaxin Wen, Zachary Ankner, Arushi Somani et al.2025 · 0 citationsarXiv
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks2601.04603v1 · Hoagy Cunningham, Jerry Wei, Zihan Wang et al.2026 · 0 citationsarXiv
Inverse Scaling in Test-Time Compute2507.14417v2 · Aryo Pradipta Gema, Alexander Hagele, Runjin Chen et al.2025 · 0 citationsTransactions on Machine Learning Research (TMLR); 12/2025; https://openreview.net/forum?id=NXgyHW1c7M
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Agentic Misalignment: How LLMs Could Be Insider Threats2510.05179v2 · Aengus Lynch, Benjamin Wright, Caleb Larson et al.2025 · 8 citationsarXivon Valency
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs2407.15549v3 · Abhay Sheshadri, Aidan Ewart, Phillip Guo et al.2024 · 3 citationsarXiv
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought2403.05518v3 · James Chua, Edward Rees, Hunar Batra et al.2024 · 0 citationsarXiv
A dataset of questions on decision-theoretic reasoning in Newcomb-like problems2411.10588v4 · Caspar Oesterheld, Emery Cooper, Miles Kodama et al.2024 · 0 citationsarXiv
Towards Understanding Sycophancy in Language Models2310.13548v4 · Mrinank Sharma, Meg Tong, Tomasz Korbak et al.2023 · 132 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Forecasting Rare Language Model Behaviors2502.16797v1 · Erik Jones, Meg Tong, Jesse Mu et al.2025 · 0 citationsarXiv
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming2501.18837v1 · Mrinank Sharma, Meg Tong, Jesse Mu et al.2025 · 7 citationsarXiv
Best-of-N Jailbreaking2412.03556v2 · John Hughes, Sara Price, Aengus Lynch et al.2024 · 0 citationsarXiv
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models2407.15211v2 · Rylan Schaeffer, Dan Valentine, Luke Bailey et al.2024 · 0 citationsarXiv
Language Models Learn to Mislead Humans via RLHF2409.12822v3 · Jiaxin Wen, Ruiqi Zhong, Akbir Khan et al.2024 · 4 citationsarXiv
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach2412.02159v1 · Tony T. Wang, John Hughes, Henry Sleight et al.2024 · 0 citationsarXiv
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats2411.17693v1 · Jiaxin Wen, Vivek Hebbar, Caleb Larson et al.2024 · 1 citationarXiv
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples2411.07494v1 · Alwin Peng, Julian Michael, Henry Sleight et al.2024 · 0 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Looking Inward: Language Models Can Learn About Themselves by Introspection2410.13787v1 · Felix J Binder, James Chua, Tomek Korbak et al.2024 · 7 citationsarXiv
Debating with More Persuasive LLMs Leads to More Truthful Answers2402.06782v4 · Akbir Khan, John Hughes, Dan Valentine et al.2024 · 11 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Inverse Scaling: When Bigger Isn't Better2306.09479v2 · Ian R. McKenzie, Alexander Lyzhov, Michael Pieler et al.2023 · 25 citationsTransactions on Machine Learning Research (TMLR), 10/2023, https://openreview.net/forum?id=DwgRm72GQF
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning2310.12921v2 · Juan Rocamonde, Victoriano Montesinos, Elvis Nava et al.2023 · 6 citationsarXiv
Improving Code Generation by Training with Natural Language Feedback2303.16749v2 · Angelica Chen, J'er'emy Scheurer, Tomasz Korbak et al.2023 · 10 citationsarXiv
Training Language Models with Language Feedback at Scale2303.16755v3 · J'er'emy Scheurer, Jon Ander Campos, Tomasz Korbak et al.2023 · 16 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting2305.04388v2 · Miles Turpin, Julian Michael, Ethan Perez et al.2023 · 196 citationsarXiv
Towards Evaluating AI Systems for Moral Status Using Self-Reports2311.08576v1 · Ethan Perez, Robert Long2023 · 31 citationsarXiv
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
Studying Large Language Model Generalization with Influence Functions2308.03296v1 · Roger Grosse, Juhan Bae, Cem Anil et al.2023 · 26 citationsarXiv
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning2307.11768v2 · Ansh Radhakrishnan, Karina Nguyen, Anna Chen et al.2023 · 8 citationsarXiv
Pretraining Language Models with Human Preferences2302.08582v2 · Tomasz Korbak, Kejian Shi, Angelica Chen et al.2023 · 25 citationsarXiv
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Training Language Models with Language Feedback2204.14146v4 · J'er'emy Scheurer, Jon Ander Campos, Jun Shern Chan et al.2022 · 11 citationsarXiv
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
RL with KL penalties is better viewed as Bayesian inference2205.11275v2 · Tomasz Korbak, Ethan Perez, Christopher L Buckley2022 · 8 citationsarXiv
Few-shot Adaptation Works with UnpredicTable Data2208.01009v2 · Jun Shern Chan, Michael Pieler, Jonathan Jao et al.2022 · 3 citationsarXiv
Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions2204.05212v2 · Alicia Parrish, Harsh Trivedi, Ethan Perez et al.2022 · 3 citationsarXiv
Red Teaming Language Models with Language Models2202.03286v1 · Ethan Perez, Saffron Huang, Francis Song et al.2022 · 347 citationsarXiv
Case-based Reasoning for Natural Language Queries over Knowledge Bases2104.08762v2 · Rajarshi Das, Manzil Zaheer, Dung Thai et al.2021 · 119 citationsarXiv
True Few-Shot Learning with Language Models2105.11447v1 · Ethan Perez, Douwe Kiela, Kyunghyun Cho2021 · 196 citationsarXiv
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks2005.11401v4 · Patrick Lewis, Ethan Perez, Aleksandra Piktus et al.2020 · 3,104 citationsarXiv
Career total: 77 works. 60 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.