SR
Sam Ringer
cs.CLcs.AIcs.LGcs.CVstat.MLcs.CYcs.NE
On Valency
published · living versionsW_63ueycek·v1 · currentpublished
Discovering Language Model Behaviors with Model-Written Evaluations
with Ethan Perez, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen +58
1 version
Preprints & journals
7 papers in the corpus · 2019–2023The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Hierarchical Quantized Autoencoders2002.08111v3 · Will Williams, Sam Ringer, Tom Ash et al.2020 · 27 citationsarXiv
Texture Bias Of CNNs Limits Few-Shot Classification Performance1910.08519v1 · Sam Ringer, Will Williams, Tom Ash et al.2019 · 6 citationsarXiv
Career total: 8 works. 7 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.