NS

Nicholas Schiefer

cs.AIcs.CLcs.LGcs.DBstat.MLcs.CRcs.CYcs.DCcs.DScs.HC

On Valency

published · living versions
W_3sjke4et·v1 · currentpublished
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
with Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert +34
1 version

Preprints & journals

18 papers in the corpus · 2019–2024
Towards Understanding Sycophancy in Language Models2310.13548v4 · Mrinank Sharma, Meg Tong, Tomasz Korbak et al.2023 · 132 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Towards Measuring the Representation of Subjective Global Opinions in Language Models2306.16388v2 · Esin Durmus, Karina Nguyen, Thomas I. Liao et al.2023 · 43 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning2307.11768v2 · Ansh Radhakrishnan, Karina Nguyen, Anna Chen et al.2023 · 8 citationsarXiv
Learned Interpolation for Better Streaming Quantile Approximation with Worst-Case Guarantees2304.07652v1 · Nicholas Schiefer, Justin Y. Chen, Piotr Indyk et al.2023 · 1 citationarXiv
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Engineering Monosemanticity in Toy Models2211.09169v1 · Adam S. Jermyn, Nicholas Schiefer, Evan Hubinger2022 · 5 citationsarXiv
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
Toy Models of Superposition2209.10652v1 · Nelson Elhage, Tristan Hume, Catherine Olsson et al.2022 · 49 citationsarXiv
FoundationDB Record Layer: A Multi-Tenant Structured Datastore1901.04452v2 · Christos Chrysafis, Ben Collins, Scott Dugas et al.2019 · 16 citationsarXiv
Career total: 29 works. 18 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.