EO
Euan Ong
cs.LGcs.AIcs.CLcs.LO
On Valency
published · living versionsW_5uruupn8·v1 · currentpublished
Auditing language models for hidden objectives
with Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey +30
1 version
Preprints & journals
4 papers in the corpus · 2023–2026Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments2608.16747v1 · Adam Karvonen, Euan Ong, Subhash Kantamneni et al.2026 · 0 citationsarXiv
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Compact Proofs of Model Performance via Mechanistic Interpretability2406.11779v14 · Jason Gross, Rajashree Agrawal, Thomas Kwa et al.2024 · 0 citationsarXiv
Successor Heads: Recurring, Interpretable Attention Heads In The Wild2312.09230v1 · Rhys Gould, Euan Ong, George Ogden et al.2023 · 3 citationsarXiv
Career total: 5 works. 4 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.