EH
Evan Hubinger
cs.AIcs.LGcs.CLcs.CRcs.CYstat.MLcs.HCcs.SE
On Valency
published · living versionsW_vv4hg27n·v1 · currentpublished
Agentic Misalignment: How LLMs Could Be Insider Threats
with Aengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie +3
1 version
Preprints & journals
19 papers in the corpus · 2019–2025Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI2511.01689v1 · Sharan Maiya, Henning Bartsch, Nathan Lambert et al.2025 · 0 citationsarXiv
Agentic Misalignment: How LLMs Could Be Insider Threats2510.05179v2 · Aengus Lynch, Benjamin Wright, Caleb Larson et al.2025 · 8 citationsarXivon Valency
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas2505.14633v1 · Yu Ying Chiu, Zhilin Wang, Sharan Maiya et al.2025 · 0 citationsarXiv
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Steering Llama 2 via Contrastive Activation Addition2312.06681v4 · Nina Panickssery, Nick Gabrieli, Julian Schulz et al.2023 · 48 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant2405.01576v1 · Olli Jarviniemi, Evan Hubinger2024 · 2 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Studying Large Language Model Generalization with Influence Functions2308.03296v1 · Roger Grosse, Juhan Bae, Cem Anil et al.2023 · 26 citationsarXiv
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning2307.11768v2 · Ansh Radhakrishnan, Karina Nguyen, Anna Chen et al.2023 · 8 citationsarXiv
Conditioning Predictive Models: Risks and Strategies2302.00805v2 · Evan Hubinger, Adam Jermyn, Johannes Treutlein et al.2023 · 0 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Engineering Monosemanticity in Toy Models2211.09169v1 · Adam S. Jermyn, Nicholas Schiefer, Evan Hubinger2022 · 5 citationsarXiv
Risks from Learned Optimization in Advanced Machine Learning Systems1906.01820v3 · Evan Hubinger, Chris van Merwijk, Vladimir Mikulik et al.2019 · 24 citationsarXiv
An overview of 11 proposals for building safe advanced AI2012.07532v1 · Evan Hubinger2020 · 6 citationsarXiv
Career total: 24 works. 19 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.