JT

Johannes Treutlein

cs.AIcs.LGcs.CLcs.GTcs.MAcs.CRecon.GNq-fin.ECstat.ML

On Valency

published · living versions
W_5uruupn8·v1 · currentpublished
Auditing language models for hidden objectives
with Samuel Marks, Trenton Bricken, Jack Lindsey, Jonathan Marcus +30
1 version

Preprints & journals

12 papers in the corpus · 2021–2026
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values2607.14345v4 · Jan Betley, Johannes Treutlein, Jan Dubi'nski et al.2026 · 0 citationsarXiv
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data2406.14546v3 · Johannes Treutlein, Dami Choi, Jan Betley et al.2024 · 5 citationsarXiv
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
Similarity-based cooperative equilibrium2211.14468v2 · Caspar Oesterheld, Johannes Treutlein, Roger Grosse et al.2022 · 1 citationarXiv
Modeling evidential cooperation in large worlds2307.04879v2 · Johannes Treutlein2023 · 0 citationsarXiv
A New Formalism, Method and Open Issues for Zero-Shot Coordination2106.06613v3 · Johannes Treutlein, Michael Dennis, Caspar Oesterheld et al.2021 · 4 citationsarXiv
Incentivizing honest performative predictions with proper scoring rules2305.17601v2 · Caspar Oesterheld, Johannes Treutlein, Emery Cooper et al.2023 · 2 citationsarXiv
Conditioning Predictive Models: Risks and Strategies2302.00805v2 · Evan Hubinger, Adam Jermyn, Johannes Treutlein et al.2023 · 0 citationsarXiv
Path Independent Equilibrium Models Can Better Exploit Test-Time Computation2211.09961v1 · Cem Anil, Ashwini Pokle, Kaiqu Liang et al.2022 · 4 citationsarXiv
COLA: Consistent Learning with Opponent-Learning Awareness2203.04098v3 · Timon Willi, Alistair Letcher, Johannes Treutlein et al.2022 · 2 citationsarXiv
Normative Disagreement as a Challenge for Cooperative AI2111.13872v1 · Julian Stastny, Maxime Rich'e, Alexander Lyzhov et al.2021 · 0 citationsarXiv
Career total: 16 works. 12 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.