DD
Dawn Drain
cs.LGcs.CLcs.SEcs.AIcs.CYcs.IRcs.PLcs.HC
On Valency
published · living versionsW_63ueycek·v1 · currentpublished
Discovering Language Model Behaviors with Model-Written Evaluations
with Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen +58
1 version
Preprints & journals
23 papers in the corpus · 2020–2023The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
In-context Learning and Induction Heads2209.11895v1 · Catherine Olsson, Nelson Elhage, Neel Nanda et al.2022 · 87 citationsarXivon Valency
Toy Models of Superposition2209.10652v1 · Nelson Elhage, Tristan Hume, Catherine Olsson et al.2022 · 49 citationsarXiv
Exploring and Evaluating Personalized Models for Code Generation2208.13928v2 · Andrei Zlotchevski, Dawn Drain, Alexey Svyatkovskiy et al.2022 · 12 citationsarXiv
Scaling Laws and Interpretability of Learning from Repeated Data2205.10487v1 · Danny Hernandez, Tom Brown, Tom Conerly et al.2022 · 22 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers2009.05634v1 · Michele Tufano, Dawn Drain, Alexey Svyatkovskiy et al.2020 · 97 citationsarXiv
A General Language Assistant as a Laboratory for Alignment2112.00861v3 · Amanda Askell, Yuntao Bai, Anna Chen et al.2021 · 27 citationsarXiv
Generating Bug-Fixes Using Pretrained Transformers2104.07896v2 · Dawn Drain, Chen Wu, Alexey Svyatkovskiy et al.2021 · 46 citationsarXiv
Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy2109.08780v1 · Colin B. Clement, Shuai Lu, Xiaoyu Liu et al.2021 · 13 citationsarXiv
GraphCodeBERT: Pre-training Code Representations with Data Flow2009.08366v4 · Daya Guo, Shuo Ren, Shuai Lu et al.2020 · 474 citationsarXiv
Distilling Transformers for Neural Cross-Domain Search2108.03322v1 · Colin B. Clement, Chen Wu, Dawn Drain et al.2021 · 1 citationarXiv
Unit Test Case Generation with Transformers and Focal Context2009.05617v2 · Michele Tufano, Dawn Drain, Alexey Svyatkovskiy et al.2020 · 8 citationsarXiv
DeepDebug: Fixing Python Bugs Using Stack Traces, Backtranslation, and Code Skeletons2105.09352v1 · Dawn Drain, Colin B. Clement, Guillermo Serrato et al.2021 · 14 citationsarXiv
Generating Code with the Help of Retrieved Template Functions and Stack Overflow Answers2104.05310v2 · Dawn Drain, Changran Hu, Chen Wu et al.2021 · 6 citationsarXiv
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation2102.04664v2 · Shuai Lu, Daya Guo, Shuo Ren et al.2021 · 443 citationsarXiv
PyMT5: multi-mode translation of natural language and Python code with transformers2010.03150v1 · Colin B. Clement, Dawn Drain, Jonathan Timcheck et al.2020 · 126 citationsarXiv
Career total: 26 works. 23 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.