SK

Saurav Kadavath

cs.CLcs.LGcs.AIcs.CVstat.MLcs.CYcs.SE

On Valency

published · living versions
W_63ueycek·v1 · currentpublished
Discovering Language Model Behaviors with Model-Written Evaluations
with Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen +58
1 version

Preprints & journals

14 papers in the corpus · 2019–2023
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
DeepChrome 2.0: Investigating and Improving Architectures, Visualizations, & Experiments2209.11923v1 · Saurav Kadavath, Samuel Paradis, Jacob Yeung2022 · 0 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
Measuring Mathematical Problem Solving With the MATH Dataset2103.03874v2 · Dan Hendrycks, Collin Burns, Saurav Kadavath et al.2021 · 276 citationsarXiv
Measuring Coding Challenge Competence With APPS2105.09938v3 · Dan Hendrycks, Steven Basart, Saurav Kadavath et al.2021 · 148 citationsarXiv
Pretraining & Reinforcement Learning: Sharpening the Axe Before Cutting the Tree2110.02497v1 · Saurav Kadavath, Samuel Paradis, Brian Yao2021 · 0 citationsarXiv
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization2006.16241v3 · Dan Hendrycks, Steven Basart, Norman Mu et al.2020 · 1,179 citationsarXiv
Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty1906.12340v2 · Dan Hendrycks, Mantas Mazeika, Saurav Kadavath et al.2019 · 560 citationsarXiv
Career total: 18 works. 14 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.