MT

Meg Tong

cs.LGcs.CLcs.AIcs.CRcs.SEstat.ML

On Valency

published · living versions
W_5uruupn8·v1 · currentpublished
Auditing language models for hidden objectives
with Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey +30
1 version

Preprints & journals

8 papers in the corpus · 2023–2025
Towards Understanding Sycophancy in Language Models2310.13548v4 · Mrinank Sharma, Meg Tong, Tomasz Korbak et al.2023 · 132 citationsarXiv
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Forecasting Rare Language Model Behaviors2502.16797v1 · Erik Jones, Meg Tong, Jesse Mu et al.2025 · 0 citationsarXiv
Steering Llama 2 via Contrastive Activation Addition2312.06681v4 · Nina Panickssery, Nick Gabrieli, Julian Schulz et al.2023 · 48 citationsarXiv
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"2309.12288v4 · Lukas Berglund, Meg Tong, Max Kaufmann et al.2023 · 31 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Taken out of context: On measuring situational awareness in LLMs2309.00667v1 · Lukas Berglund, Asa Cooper Stickland, Mikita Balesni et al.2023 · 8 citationsarXiv
Career total: 10 works. 8 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.