OE
Owain Evans
cs.AIcs.LGcs.CLcs.CRcs.CYstat.MLcs.CVcs.NEepidemiologyLarge Language Models
On Valency
published · living versionsW_y3274kcu·v1 · currentpublished
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
with Miles Brundage, Shahar Avin, Jack Clark, Helen Toner +21
1 version
Preprints & journals
39 papers in the corpus · 2015–2026Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble2609.10883v1 · Jorio Cocola, Lev McKinney, Harry Mayne et al.2026 · 0 citationsarXiv
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values2607.14345v4 · Jan Betley, Johannes Treutlein, Jan Dubi'nski et al.2026 · 0 citationsarXiv
Negation Neglect: When models fail to learn negations in training2605.13829v1 · Harry Mayne, Lev McKinney, Jan Dubi'nski et al.2026 · 0 citationsarXiv
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers2604.25891v1 · Jan Dubi'nski, Jan Betley, Anna Sztyber-Betley et al.2026 · 0 citationsarXiv
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious2604.13051v1 · James Chua, Jan Betley, Samuel Marks et al.2026 · 0 citationsarXiv
Language models transmit behavioural traits through hidden signals in data.41986627 · Cloud, Alex, Le, Minh, Chua, James et al.2026 · 9 citationsNature. 2026;652(8110):615-621
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs2502.17424v7 · Jan Betley, Daniel Tan, Niels Warncke et al.2025 · 27 citationsPMLR 267:4043-4068
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs2512.09742v1 · Jan Betley, Jorio Cocola, Dylan Feng et al.2025 · 2 citationsarXiv
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Lessons from Studying Two-Hop Latent Reasoning2411.16353v4 · Mikita Balesni, Tomek Korbak, Owain Evans2024 · 0 citationsarXiv
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data2507.14805v1 · Alex Cloud, Minh Le, James Chua et al.2025 · 3 citationsarXiv
Are DeepSeek R1 And Other Reasoning Models More Faithful?2501.08156v5 · James Chua, Owain Evans2025 · 0 citationsarXiv
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models2506.13206v2 · James Chua, Jan Betley, Mia Taylor et al.2025 · 1 citationarXiv
Tell me about yourself: LLMs are aware of their learned behaviors2501.11120v1 · Jan Betley, Xuchan Bao, Mart'in Soto et al.2025 · 2 citationsarXiv
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data2406.14546v3 · Johannes Treutlein, Dami Choi, Jan Betley et al.2024 · 5 citationsarXiv
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation1802.07228v2 · Miles Brundage, Shahar Avin, Jack Clark et al.2018 · 495 citationsarXivon Valency
Towards evaluations-based safety cases for AI scheming2411.03336v2 · Mikita Balesni, Marius Hobbhahn, David Lindner et al.2024 · 2 citationsarXiv
Looking Inward: Language Models Can Learn About Themselves by Introspection2410.13787v1 · Felix J Binder, James Chua, Tomek Korbak et al.2024 · 7 citationsarXiv
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs2407.04694v1 · Rudolf Laine, Bilal Chughtai, Jan Betley et al.2024 · 7 citationsarXiv
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"2309.12288v4 · Lukas Berglund, Meg Tong, Max Kaufmann et al.2023 · 31 citationsarXiv
Can Language Models Explain Their Own Classification Behavior?2405.07436v1 · Dane Sherburn, Bilal Chughtai, Owain Evans2024 · 0 citationsarXiv
Tell, don't show: Declarative facts influence how LLMs generalize2312.07779v1 · Alexander Meinke, Owain Evans2023 · 0 citationsarXiv
How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions2309.15840v1 · Lorenzo Pacchiardi, Alex J. Chan, Soren Mindermann et al.2023 · 7 citationsarXiv
Taken out of context: On measuring situational awareness in LLMs2309.00667v1 · Lukas Berglund, Asa Cooper Stickland, Mikita Balesni et al.2023 · 8 citationsarXiv
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models2206.04615v3 · Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al.2022 · 565 citationsTransactions on Machine Learning Research, May/2022, https://openreview.net/forum?id=uyTL5Bvosj
Forecasting Future World Events with Neural Networks2206.15474v2 · Andy Zou, Tristan Xiao, Ryan Jia et al.2022 · 7 citationsarXiv
Teaching Models to Express Their Uncertainty in Words2205.14334v2 · Stephanie Lin, Jacob Hilton, Owain Evans2022 · 54 citationsarXiv
TruthfulQA: Measuring How Models Mimic Human Falsehoods2109.07958v2 · Stephanie Lin, Jacob Hilton, Owain Evans2021 · 825 citationsarXiv
Truthful AI: Developing and governing AI that does not lie2110.06674v1 · Owain Evans, Owen Cotton-Barratt, Lukas Finnveden et al.2021 · 9 citationsarXiv
Active Reinforcement Learning: Observing Rewards at a Cost2011.06709v2 · David Krueger, Jan Leike, Owain Evans et al.2020 · 12 citationsarXiv
Estimating Household Transmission of SARS-CoV-210.1101/2020.05.23.20111559v2 · Curmei, M., Ilyas, A., Evans, O. et al.2020 · 32 citationsmedRxiv
Sensory Optimization: Neural Networks as a Model for Understanding and Creating Art1911.07068v1 · Owain Evans2019 · 1 citationarXiv
Generalizing from a few environments in safety-critical reinforcement learning1907.01475v1 · Zachary Kenton, Angelos Filos, Owain Evans et al.2019 · 8 citationsarXiv
When Will AI Exceed Human Performance? Evidence from AI Experts1705.08807v3 · Katja Grace, John Salvatier, Allan Dafoe et al.2017 · 222 citationsarXivon Valency
Active Reinforcement Learning with Monte-Carlo Tree Search1803.04926v3 · Sebastian Schulze, Owain Evans2018 · 6 citationsarXiv
Trial without Error: Towards Safe Reinforcement Learning via Human Intervention1707.05173v1 · William Saunders, Girish Sastry, Andreas Stuhlmueller et al.2017 · 170 citationsarXiv
Agent-Agnostic Human-in-the-Loop Reinforcement Learning1701.04079v1 · David Abel, John Salvatier, Andreas Stuhlmuller et al.2017 · 22 citationsarXiv
Learning the Preferences of Ignorant, Inconsistent Agents1512.05832v1 · Owain Evans, Andreas Stuhlmueller, Noah D. Goodman2015 · 96 citationsarXiv
Career total: 52 works. 39 are in this corpus.Showing the 38 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.