JH

Johannes Heidecke

cs.AIcs.CLcs.LGcs.CYcs.CRcs.CVcs.SDcs.SEeess.AS

On Valency

published · living versions
W_jsb5733n·v1 · currentpublished
Deliberative Alignment: Reasoning Enables Safer Language Models
with Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain +10
1 version

Preprints & journals

15 papers in the corpus · 2022–2026
Reinforcement Learning Towards Broadly and Persistently Beneficial Models2606.24014v1 · Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab et al.2026 · 0 citationsarXiv
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats2604.27470v1 · Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim et al.2026 · 1 citationarXiv
Persona Features Control Emergent Misalignment2506.19823v2 · Miles Wang, Tom Dupr'e la Tour, Olivia Watkins et al.2025 · 2 citationsarXivon Valency
AI-based Clinical Decision Support for Primary Care: A Real-World Study2507.16947v1 · Robert Korom, Sarah Kiptinness, Najib Adan et al.2025 · 10 citationsarXiv
The Singapore Consensus on Global AI Safety Research Priorities2506.20702v2 · Yoshua Bengio, Tegan Maharaj, Luke Ong et al.2025 · 5 citationsarXiv
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?2502.12115v4 · Samuel Miserendino, Michele Wang, Tejal Patwardhan et al.2025 · 4 citationsarXiv
First-Person Fairness in Chatbots2410.19803v2 · Tyna Eloundou, Alex Beutel, David G. Robinson et al.2024 · 3 citationsarXiv
Trading Inference-Time Compute for Adversarial Robustness2501.18841v1 · Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak et al.2025 · 0 citationsarXivon Valency
Deliberative Alignment: Reasoning Enables Safer Language Models2412.16339v2 · Melody Y. Guan, Manas Joglekar, Eric Wallace et al.2024 · 34 citationsarXivon Valency
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning2412.18693v1 · Alex Beutel, Kai Xiao, Johannes Heidecke et al.2024 · 1 citationarXiv
Rule Based Rewards for Language Model Safety2411.01111v1 · Tong Mu, Alec Helyar, Johannes Heidecke et al.2024 · 10 citationsarXiv
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions2404.13208v1 · Eric Wallace, Kai Xiao, Reimar Leike et al.2024 · 10 citationsarXiv
Text and Code Embeddings by Contrastive Pre-Training2201.10005v1 · Arvind Neelakantan, Tao Xu, Raul Puri et al.2022 · 149 citationsarXiv
Career total: 19 works. 15 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.