SM
Samuel Marks
cs.AIcs.CLcs.LG
On Valency
published · living versionsW_8zqhdax4·v1 · currentpublished
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
with Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud +9
1 version
Preprints & journals
3 papers in the corpus · 2024–2026Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation2603.05494v2 · Helena Casademunt, Bartosz Cywi'nski, Khoi Tran et al.2026 · 0 citationsarXiv
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data2406.14546v3 · Johannes Treutlein, Dami Choi, Jan Betley et al.2024 · 5 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Career total: 5 works. 3 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.