AR

Ansh Radhakrishnan

cs.CLcs.AIcs.LGcs.CRcs.SE

On Valency

published · living versions
W_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Jonathan Uesato, Carson Denison +10
1 version

Preprints & journals

6 papers in the corpus · 2023–2025
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats2411.17693v1 · Jiaxin Wen, Vivek Hebbar, Caleb Larson et al.2024 · 1 citationarXiv
Debating with More Persuasive LLMs Leads to More Truthful Answers2402.06782v4 · Akbir Khan, John Hughes, Dan Valentine et al.2024 · 11 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning2307.11768v2 · Ansh Radhakrishnan, Karina Nguyen, Anna Chen et al.2023 · 8 citationsarXiv
Career total: 7 works. 6 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.