RG
Ryan Greenblatt
cs.LGcs.AIcs.CLcs.CRcs.SEstat.ML
On Valency
published · living versionsW_uutna755·v1 · currentpublished
Alignment faking in large language models
with Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid +15
1 version
Preprints & journals
9 papers in the corpus · 2023–2026Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models2606.07157v4 · Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff et al.2026 · 0 citationsarXiv
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?2510.17941v1 · Stewart Slocum, Julian Minder, Cl'ement Dumas et al.2025 · 0 citationsarXiv
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
AI Control: Improving Safety Despite Intentional Subversion2312.06942v5 · Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan et al.2023 · 3 citationsarXiv
Stress-Testing Capability Elicitation With Password-Locked Models2405.19550v1 · Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov et al.2024 · 0 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Preventing Language Models From Hiding Their Reasoning2310.18512v2 · Fabien Roger, Ryan Greenblatt2023 · 2 citationsarXiv
Benchmarks for Detecting Measurement Tampering2308.15605v5 · Fabien Roger, Ryan Greenblatt, Max Nadeau et al.2023 · 0 citationsarXiv
Career total: 13 works. 9 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.