DF

David Farhi

cs.LGstat.MLcs.AIcs.CLcs.CVcs.CYcs.SDeess.AS

On Valency

published · living versions
W_fvh7g6bt·v1 · currentpublished
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
with Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou +4
1 version

Preprints & journals

3 papers in the corpus · 2019–2024
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft2106.14876v1 · Ingmar Kanitscheider, Joost Huizinga, David Farhi et al.2021 · 9 citationsarXiv
Dota 2 with Large Scale Deep Reinforcement Learning1912.06680v1 · OpenAI: Christopher Berner, Greg Brockman, Brooke Chan et al.2019 · 1,036 citationsarXiv
Career total: 3 works. 3 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.