KS
Kshitij Sachan
cs.AIcs.LGcs.CLcs.CRcs.NEcs.SE
On Valency
published · living versionsW_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version
Preprints & journals
4 papers in the corpus · 2022–2024Polysemanticity and Capacity in Neural Networks2210.01892v4 · Adam Scherlis, Kshitij Sachan, Adam S. Jermyn et al.2022 · 5 citationsarXiv
Debating with More Persuasive LLMs Leads to More Truthful Answers2402.06782v4 · Akbir Khan, John Hughes, Dan Valentine et al.2024 · 11 citationsarXiv
AI Control: Improving Safety Despite Intentional Subversion2312.06942v5 · Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan et al.2023 · 3 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Career total: 4 works. 4 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.