cs.LGcs.AIcs.CLstat.MLcs.CRcs.CVcs.CYcs.NE

On Valency

published · living versions
W_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison +10
1 version

Preprints & journals

19 papers in the corpus · 2022–2026
Diffuse AI Control on Fuzzy Tasks2606.08892v2 · Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar et al.2026 · 0 citationsarXiv
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors2605.16626v2 · Elle Najt, Colin Toft, Tyler Tracy et al.2026 · 0 citationsarXiv
Removing Sandbagging in LLMs by Training with Weak Supervision2604.22082v2 · Emil Ryd, Henning Bartsch, Julian Stastny et al.2026 · 0 citationsarXiv
Inverse Scaling in Test-Time Compute2507.14417v2 · Aryo Pradipta Gema, Alexander Hagele, Runjin Chen et al.2025 · 0 citationsTransactions on Machine Learning Research (TMLR); 12/2025; https://openreview.net/forum?id=NXgyHW1c7M
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning2506.22777v2 · Miles Turpin, Andy Arditi, Marvin Li et al.2025 · 0 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Polysemanticity and Capacity in Neural Networks2210.01892v4 · Adam Scherlis, Kshitij Sachan, Adam S. Jermyn et al.2022 · 5 citationsarXiv
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models2407.15211v2 · Rylan Schaeffer, Dan Valentine, Luke Bailey et al.2024 · 0 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Nearly $d$-Linear Convergence Bounds for Diffusion Models via Stochastic Localization2308.03686v3 · Joe Benton, Valentin De Bortoli, Arnaud Doucet et al.2023 · 1 citationarXiv
From Denoising Diffusions to Denoising Markov Models2211.03595v3 · Joe Benton, Yuyang Shi, Valentin De Bortoli et al.2022 · 9 citationsarXiv
Error Bounds for Flow Matching Methods2305.16860v2 · Joe Benton, George Deligiannidis, Arnaud Doucet2023 · 3 citationsarXiv
Measuring Feature Sparsity in Language Models2310.07837v2 · Mingyang Deng, Lucas Tao, Joe Benton2023 · 1 citationarXiv
A Continuous Time Framework for Discrete Denoising Models2205.14987v2 · Andrew Campbell, Joe Benton, Valentin De Bortoli et al.2022 · 16 citationsarXiv
Career total: 23 works. 19 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.