LG

Logan Graham

cs.AIcs.LGcs.CLcs.CRcs.CYcs.PLcs.SEstat.COstat.ML

On Valency

published · living versions
W_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version

Preprints & journals

4 papers in the corpus · 2019–2025
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims2004.07213v2 · Miles Brundage, Shahar Avin, Jasmine Wang et al.2020 · 301 citationsarXiv
MultiVerse: Causal Reasoning using Importance Sampling in Probabilistic Programming1910.08091v2 · Yura Perov, Logan Graham, Kostis Gourgoulias et al.2019 · 13 citationsarXiv
Career total: 7 works. 4 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.