JU

Jonathan Uesato

cs.LGcs.AIcs.CLstat.MLcs.CRcs.CYcs.CVcs.NEcs.PLcs.SD

On Valency

published · living versions
W_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Ansh Radhakrishnan, Carson Denison +10
1 version

Preprints & journals

30 papers in the corpus · 2016–2025
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
Gemini: A Family of Highly Capable Multimodal Models2312.11805v5 · Gemini Team Google: Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et al.2023 · 838 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Solving math word problems with process- and outcome-based feedback2211.14275v1 · Jonathan Uesato, Nate Kushman, Ramana Kumar et al.2022 · 23 citationsarXiv
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals2210.01790v2 · Rohin Shah, Vikrant Varma, Ramana Kumar et al.2022 · 15 citationsarXiv
Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models2206.08325v2 · Maribeth Rauh, John Mellor, Jonathan Uesato et al.2022 · 14 citationsarXiv
Improving alignment of dialogue agents via targeted human judgements2209.14375v1 · Amelia Glaese, Nat McAleese, Maja Trebacz et al.2022 · 132 citationsarXiv
Scaling Language Models: Methods, Analysis & Insights from Training Gopher2112.11446v2 · Jack W. Rae, Sebastian Borgeaud, Trevor Cai et al.2021 · 245 citationsarXiv
Ethical and social risks of harm from Language Models2112.04359v1 · Laura Weidinger, John Mellor, Maribeth Rauh et al.2021 · 71 citationsarXiv
Make Sure You're Unsure: A Framework for Verifying Probabilistic Specifications2102.09479v2 · Leonard Berrada, Sumanth Dathathri, Krishnamurthy Dvijotham et al.2021 · 6 citationsarXiv
An Empirical Investigation of Learning from Biased Toxicity Labels2110.01577v1 · Neel Nanda, Jonathan Uesato, Sven Gowal2021 · 0 citationsarXiv
Challenges in Detoxifying Language Models2109.07445v1 · Johannes Welbl, Amelia Glaese, Jonathan Uesato et al.2021 · 102 citationsarXiv
Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples2010.03593v3 · Sven Gowal, Chongli Qin, Jonathan Uesato et al.2020 · 139 citationsarXiv
REALab: An Embedded Perspective on Tampering2011.08820v1 · Ramana Kumar, Jonathan Uesato, Richard Ngo et al.2020 · 3 citationsarXiv
Avoiding Tampering Incentives in Deep RL via Decoupled Approval2011.08827v1 · Jonathan Uesato, Ramana Kumar, Victoria Krakovna et al.2020 · 2 citationsarXiv
Enabling certification of verification-agnostic networks via memory-efficient semidefinite programming2010.11645v2 · Sumanth Dathathri, Krishnamurthy Dvijotham, Alexey Kurakin et al.2020 · 28 citationsarXiv
Are Labels Required for Improving Adversarial Robustness?1905.13725v4 · Jonathan Uesato, Jean-Baptiste Alayrac, Po-Sen Huang et al.2019 · 192 citationsarXiv
An Alternative Surrogate Loss for PGD-based Adversarial Testing1910.09338v1 · Sven Gowal, Jonathan Uesato, Chongli Qin et al.2019 · 46 citationsarXiv
On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models1810.12715v4 · Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth et al.2018 · 280 citationsarXiv
Verification of Non-Linear Specifications for Neural Networks1902.09592v1 · Chongli Qin, Krishnamurthy (Dj) Dvijotham, Brendan O'Donoghue et al.2019 · 22 citationsarXiv
Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures1812.01647v1 · Jonathan Uesato, Ananya Kumar, Csaba Szepesvari et al.2018 · 61 citationsarXiv
Robustness via curvature regularization, and vice versa1811.09716v1 · Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Jonathan Uesato et al.2018 · 240 citationsarXiv
Strength in Numbers: Trading-off Robustness and Computation via Adversarially-Trained Ensembles1811.09300v1 · Edward Grefenstette, Robert Stanforth, Brendan O'Donoghue et al.2018 · 9 citationsarXiv
Technical Report on the CleverHans v2.1.0 Adversarial Examples Library1610.00768v6 · Nicolas Papernot, Fartash Faghri, Nicholas Carlini et al.2016 · 401 citationsarXiv
Adversarial Risk and the Dangers of Evaluating Against Weak Attacks1802.05666v2 · Jonathan Uesato, Brendan O'Donoghue, Aaron van den Oord et al.2018 · 408 citationsarXiv
Training verified learners with learned verifiers1805.10265v2 · Krishnamurthy Dvijotham, Sven Gowal, Robert Stanforth et al.2018 · 87 citationsarXiv
Semantic Code Repair using Neuro-Symbolic Transformation Networks1710.11054v1 · Jacob Devlin, Jonathan Uesato, Rishabh Singh et al.2017 · 27 citationsarXiv
RobustFill: Neural Program Learning under Noisy I/O1703.07469v1 · Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju et al.2017 · 95 citationsarXiv
Career total: 37 works. 30 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.