PH
Peter Hase
cs.CLcs.AIcs.LGcs.CVcs.CYcs.SEeess.IVstat.ML
On Valency
published · living versionsW_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato +10
1 version
Preprints & journals
27 papers in the corpus · 2018–2026Comparing Linear Probes with Mahalanobis Cosine Similarity2606.19603v1 · Zhuofan Josh Ying, Peter Hase, Nikolaus Kriegeskorte2026 · 0 citationsarXiv
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning2604.22074v1 · Qinan Yu, Alexa Tartaglini, Peter Hase et al.2026 · 0 citationsarXiv
Unsupervised Elicitation of Language Models2506.10139v2 · Jiaxin Wen, Zachary Ankner, Arushi Somani et al.2025 · 0 citationsarXiv
Are language models rational? The case of coherence norms and belief revision2406.03442v3 · Thomas Hofweber, Peter Hase, Elias Stengel-Eskin et al.2024 · 0 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation2505.01456v1 · Vaidehi Patil, Yi-Lin Sung, Peter Hase et al.2025 · 0 citationsarXiv
System-1.x: Learning to Balance Fast and Slow Planning with Language Models2407.14414v2 · Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen et al.2024 · 1 citationarXiv
Teaching Models to Balance Resisting and Accepting Persuasion2410.14596v2 · Elias Stengel-Eskin, Peter Hase, Mohit Bansal2024 · 0 citationsarXiv
Foundational Challenges in Assuring Alignment and Safety of Large Language Models2404.09932v2 · Usman Anwar, Abulhair Saparov, Javier Rando et al.2024 · 16 citationsarXiv
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models2405.21028v2 · Elias Stengel-Eskin, Peter Hase, Mohit Bansal2024 · 1 citationarXiv
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?2406.19354v1 · Peter Hase, Thomas Hofweber, Xiang Zhou et al.2024 · 1 citationarXiv
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks2401.06751v2 · Peter Hase, Mohit Bansal, Peter Clark et al.2024 · 3 citationsarXiv
Can Language Models Teach Weaker Agents? Teacher Explanations Improve Students via Personalization2306.09299v2 · Swarnadeep Saha, Peter Hase, Mohit Bansal2023 · 1 citationarXiv
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models2301.04213v2 · Peter Hase, Mohit Bansal, Been Kim et al.2023 · 30 citationsarXiv
Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks2309.17410v1 · Vaidehi Patil, Peter Hase, Mohit Bansal2023 · 4 citationsarXiv
GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models2203.07281v2 · Archiki Prasad, Peter Hase, Xiang Zhou et al.2022 · 76 citationsarXiv
Summarization Programs: Interpretable Abstractive Summarization with Neural Modular Trees2209.10492v2 · Swarnadeep Saha, Shiyue Zhang, Peter Hase et al.2022 · 7 citationsarXiv
Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations2211.07517v1 · Swarnadeep Saha, Peter Hase, Nazneen Rajani et al.2022 · 8 citationsarXiv
Low-Cost Algorithmic Recourse for Users With Uncertain Cost Functions2111.01235v2 · Prateek Yadav, Peter Hase, Mohit Bansal2021 · 0 citationsarXiv
Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs2111.13654v1 · Peter Hase, Mona Diab, Asli Celikyilmaz et al.2021 · 13 citationsarXiv
The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations2106.00786v2 · Peter Hase, Harry Xie, Mohit Bansal2021 · 24 citationsarXiv
FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging2012.15781v2 · Han Guo, Nazneen Fatema Rajani, Peter Hase et al.2020 · 56 citationsarXiv
When Can Models Learn From Explanations? A Formal Framework for Understanding the Roles of Explanation Data2102.02201v2 · Peter Hase, Mohit Bansal2021 · 41 citationsarXiv
Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language?2010.04119v1 · Peter Hase, Shiyue Zhang, Harry Xie et al.2020 · 62 citationsarXiv
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?2005.01831v1 · Peter Hase, Mohit Bansal2020 · 254 citationsarXiv
Shall I Compare Thee to a Machine-Written Sonnet? An Approach to Algorithmic Sonnet Generation1811.05067v2 · John Benhardt, Peter Hase, Liuyi Zhu et al.2018 · 5 citationsarXiv
Interpretable Image Recognition with Hierarchical Prototypes1906.10651v2 · Peter Hase, Chaofan Chen, Oscar Li et al.2019 · 128 citationsarXiv
Career total: 45 works. 27 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.