cs.AIcs.LGcs.CLcs.CYcs.CRcs.CVcs.DBcs.IRcs.ROcs.SE

On Valency

published · living versions
W_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version

Preprints & journals

59 papers in the corpus · 2023–2026
Running the Gauntlet: Hard Agentic Tasks2606.14397v5 · Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel et al.2026 · 0 citationsarXiv
The 2026 Singapore Consensus on Global AI Safety Research Priorities2608.14611v1 · Stephen Casper, Oskar Galeev, Yoshua Bengio et al.2026 · 0 citationsarXiv
Understanding Addition and Subtraction in Transformers2402.02619v11 · Philip Quirke, Clement Neo, Fazl Barez2024 · 0 citationsarXiv
AIMO Interpretability Challenge2607.13899v1 · Michal Stef'anik, Philipp Mondorf, Andreas Waldis et al.2026 · 0 citationsarXiv
Pretraining Curricula Enable Selective Fine-tuning2607.04846v1 · Sebastian A. Bruijns, Jirko Rubruck, Mia H. Whitefield et al.2026 · 0 citationsarXiv
From Democracies to Autocracies: How AI Systems Enable Authoritarianism by Design2606.17286v1 · Jeba Sania, Marta Ziosi, Fazl Barez2026 · 0 citationsarXiv
Position: Token Taxes Can Mitigate AI's Economic Risks2603.04555v2 · Lucas Irwin, Tung-Yu Wu, Fazl Barez2026 · 0 citationsarXiv
Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics2606.06533v1 · Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah et al.2026 · 0 citationsarXiv
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing2606.00033v1 · Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi et al.2026 · 0 citationsarXiv
Chain-of-Thought Hijacking2510.26418v4 · Jianli Zhao, Tingchen Fu, Rylan Schaeffer et al.2025 · 0 citationsarXiv
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs2603.03308v2 · Adi Simhi, Fazl Barez, Martin Tutek et al.2026 · 0 citationsarXiv
Interpretability Can Be Actionable2605.11161v1 · Hadas Orgad, Fazl Barez, Tal Haklay et al.2026 · 0 citationsarXiv
Rigorous Interpretation Is a Form of Evaluation2605.05508v1 · Isabelle Lee, Emmy Liu, Cathy Jiao et al.2026 · 0 citationsarXiv
Beyond alignment: Why robotic foundation models need context-aware safety.42054474 · Robey, Alexander, Ravichandran, Zachary, Jones, Eliot Krzysztof et al.2026 · 0 citationsScience robotics. 2026;11(113):eaef2191
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models2509.26238v4 · James Oldfield, Philip Torr, Ioannis Patras et al.2025 · 0 citationsarXiv
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models2505.24535v3 · Narmeen Oozeer, Luke Marks, Shreyans Jain et al.2025 · 0 citationsarXiv
Curveball Steering: The Right Direction To Steer Isn't Always Linear2603.09313v3 · Shivam Raval, Hae Jin Song, Linlin Wu et al.2026 · 0 citationsarXiv
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models2510.05465v3 · Aman Gupta, Denny O'Shea, Fazl Barez2025 · 0 citationsarXiv
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value2512.03399v1 · Joe Edelman, Tan Zhi-Xuan, Ryan Lowe et al.2025 · 0 citationsarXiv
Precise In-Parameter Concept Erasure in Large Language Models2505.22586v2 · Yoav Gur-Arieh, Clara Suslik, Yihuai Hong et al.2025 · 1 citationarXiv
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models2506.00051v2 · Philip Quirke, Narmeen Oozeer, Chaithanya Bandi et al.2025 · 0 citationsarXiv
Interpreting Learned Feedback Patterns in Large Language Models2310.08164v6 · Luke Marks, Amir Abdullah, Clement Neo et al.2023 · 0 citationsarXiv
Do Sparse Autoencoders Generalize? A Case Study of Answerability2502.19964v2 · Lovis Heindrich, Philip Torr, Fazl Barez et al.2025 · 0 citationsarXiv
Embodied AI: Emerging Risks and Opportunities for Policy Action2509.00117v2 · Jared Perlo, Alexander Robey, Fazl Barez et al.2025 · 1 citationarXiv
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer2502.12964v2 · Adi Simhi, Itay Itzhak, Fazl Barez et al.2025 · 7 citationsarXiv
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective2508.12531v1 · Minseon Kim, Jin Myung Kwak, Lama Alssum et al.2025 · 0 citationsarXiv
The Singapore Consensus on Global AI Safety Research Priorities2506.20702v2 · Yoshua Bengio, Tegan Maharaj, Luke Ong et al.2025 · 5 citationsarXiv
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning2410.08811v2 · Tingchen Fu, Mrinank Sharma, Philip Torr et al.2024 · 2 citationsarXiv
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders2410.06981v4 · Michael Lan, Philip Torr, Austin Meek et al.2024 · 2 citationsarXiv
Towards Interpreting Visual Information Processing in Vision-Language Models2410.07149v2 · Clement Neo, Luke Ong, Philip Torr et al.2024 · 3 citationsarXiv
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?2504.12914v1 · Ben Bucknall, Saad Siddiqui, Lara Thurnherr et al.2025 · 2 citationsarXiv
Rethinking AI Cultural Alignment2501.07751v2 · Michal Bravansky, Filip Trhlik, Fazl Barez2025 · 1 citationarXiv
Open Problems in Machine Unlearning for AI Safety2501.04952v1 · Fazl Barez, Tingchen Fu, Ameya Prabhu et al.2025 · 0 citationsarXiv
Best-of-N Jailbreaking2412.03556v2 · John Hughes, Sara Price, Aengus Lynch et al.2024 · 0 citationsarXiv
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders2411.01220v2 · Luke Marks, Alasdair Paren, David Krueger et al.2024 · 4 citationsarXiv
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions2402.15055v2 · Clement Neo, Shay B. Cohen, Fazl Barez2024 · 3 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Risks and Opportunities of Open-Source Generative AI2405.08597v3 · Francisco Eiras, Aleksandar Petrov, Bertie Vidgen et al.2024 · 8 citationsarXiv
Near to Mid-term Risks and Opportunities of Open-Source Generative AI2404.17047v2 · Francisco Eiras, Aleksandar Petrov, Bertie Vidgen et al.2024 · 3 citationsarXiv
Visualizing Neural Network Imagination2405.06409v1 · Nevan Wichers, Victor Tao, Riccardo Volpato et al.2024 · 0 citationsarXiv
Understanding Addition in Transformers2310.13121v9 · Philip Quirke, Fazl Barez2023 · 0 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Large Language Models Relearn Removed Concepts2401.01814v1 · Michelle Lo, Shay B. Cohen, Fazl Barez2024 · 4 citationsarXiv
Measuring Value Alignment2312.15241v1 · Fazl Barez, Philip Torr2023 · 0 citationsNeurIPS 2023 MP2 Workshop
DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models2310.01870v2 · Albert Garde, Esben Kran, Fazl Barez2023 · 2 citationsarXiv
Career total: 90 works. 59 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.