CA
Cem Anil
cs.LGcs.AIcs.CLstat.MLcs.CRcs.CVcs.CYcs.GTcs.SDcs.SE
On Valency
published · living versionsW_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version
Preprints & journals
16 papers in the corpus · 2018–2026Modular Pretraining Enables Access Control2607.08077v1 · Ethan Roland, Murat Cubuktepe, Erick Martinez et al.2026 · 0 citationsarXiv
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs2512.05648v1 · Igor Shilov, Alex Cloud, Aryo Pradipta Gema et al.2025 · 0 citationsarXiv
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming2501.18837v1 · Mrinank Sharma, Meg Tong, Jesse Mu et al.2025 · 7 citationsarXiv
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data2406.14546v3 · Johannes Treutlein, Dami Choi, Jan Betley et al.2024 · 5 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) Pipeline for Musical Timbre Transfer1811.09620v3 · Sicong Huang, Qiyang Li, Cem Anil et al.2018 · 75 citationsICLR 2019
Studying Large Language Model Generalization with Influence Functions2308.03296v1 · Roger Grosse, Juhan Bae, Cem Anil et al.2023 · 26 citationsarXiv
Path Independent Equilibrium Models Can Better Exploit Test-Time Computation2211.09961v1 · Cem Anil, Ashwini Pokle, Kaiqu Liang et al.2022 · 4 citationsarXiv
Exploring Length Generalization in Large Language Models2207.04901v2 · Cem Anil, Yuhuai Wu, Anders Andreassen et al.2022 · 49 citationsarXiv
Solving Quantitative Reasoning Problems with Language Models2206.14858v2 · Aitor Lewkowycz, Anders Andreassen, David Dohan et al.2022 · 339 citationsarXiv
Learning to Elect2108.02768v3 · Cem Anil, Xuchan Bao2021 · 0 citationsarXiv
Learning to Give Checkable Answers with Prover-Verifier Games2108.12099v1 · Cem Anil, Guodong Zhang, Yuhuai Wu et al.2021 · 0 citationsarXiv
Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks1911.00937v2 · Qiyang Li, Saminul Haque, Cem Anil et al.2019 · 40 citationsarXiv
Sorting out Lipschitz function approximation1811.05381v2 · Cem Anil, James Lucas, Roger Grosse2018 · 70 citationsarXiv
Training Deep Networks with Synthetic Data: Bridging the Reality Gap by Domain Randomization1804.06516v3 · Jonathan Tremblay, Aayush Prakash, David Acuna et al.2018 · 914 citationsarXiv
Career total: 23 works. 16 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.