AH

Alec Helyar

cs.AIcs.CLcs.CYcs.LGcs.CR

On Valency

published · living versions
W_jsb5733n·v1 · currentpublished
Deliberative Alignment: Reasoning Enables Safer Language Models
with Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain +10
1 version

Preprints & journals

6 papers in the corpus · 2023–2025
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training2508.09224v1 · Yuan Yuan, Tina Sriskandarajah, Anna-Luisa Brakman et al.2025 · 5 citationsarXiv
Deliberative Alignment: Reasoning Enables Safer Language Models2412.16339v2 · Melody Y. Guan, Manas Joglekar, Eric Wallace et al.2024 · 34 citationsarXivon Valency
Rule Based Rewards for Language Model Safety2411.01111v1 · Tong Mu, Alec Helyar, Johannes Heidecke et al.2024 · 10 citationsarXiv
A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications2310.17750v1 · Ahmed Magooda, Alec Helyar, Kyle Jackson et al.2023 · 2 citationsarXiv
Transformer-based Vulnerability Detection in Code at EditTime: Zero-shot, Few-shot, or Fine-tuning?2306.01754v1 · Aaron Chan, Anant Kharkar, Roshanak Zilouchian Moghaddam et al.2023 · 7 citationsarXiv
Career total: 8 works. 6 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.