CO

Catherine Olsson

cs.LGcs.CLcs.AIstat.MLcs.CYcs.CRcs.CVLeisure ActivitiesSocial Participation

On Valency

published · living versions
W_63ueycek·v1 · currentpublished
Discovering Language Model Behaviors with Model-Written Evaluations
with Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen +58
1 version

Preprints & journals

19 papers in the corpus · 2011–2025
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
In-context Learning and Induction Heads2209.11895v1 · Catherine Olsson, Nelson Elhage, Neel Nanda et al.2022 · 87 citationsarXivon Valency
Toy Models of Superposition2209.10652v1 · Nelson Elhage, Tristan Hume, Catherine Olsson et al.2022 · 49 citationsarXiv
Scaling Laws and Interpretability of Learning from Repeated Data2205.10487v1 · Danny Hernandez, Tom Brown, Tom Conerly et al.2022 · 22 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
A General Language Assistant as a Laboratory for Alignment2112.00861v3 · Amanda Askell, Yuntao Bai, Anna Chen et al.2021 · 27 citationsarXiv
Dota 2 with Large Scale Deep Reinforcement Learning1912.06680v1 · OpenAI: Christopher Berner, Greg Brockman, Brooke Chan et al.2019 · 1,036 citationsarXiv
Discriminator Rejection Sampling1810.06758v3 · Samaneh Azadi, Catherine Olsson, Trevor Darrell et al.2018 · 26 citationsarXiv
Unrestricted Adversarial Examples1809.08352v1 · Tom B. Brown, Nicholas Carlini, Chiyuan Zhang et al.2018 · 70 citationsarXiv
Skill Rating for Generative Models1808.04888v1 · Catherine Olsson, Surya Bhupatiraju, Tom Brown et al.2018 · 23 citationsarXiv
Is Generator Conditioning Causally Related to GAN Performance?1802.08768v2 · Augustus Odena, Jacob Buckman, Catherine Olsson et al.2018 · 48 citationsarXiv
Activity participation of children with complex communication needs, physical disabilities and typically-developing peers.21548855 · Raghavendra, Parimala, Virgo, Rachael, Olsson, Catherine et al.2011 · 74 citationsDevelopmental neurorehabilitation. 2011;14(3):145-55
Career total: 21 works. 19 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.