AT
Alex Tamkin
cs.CLcs.LGcs.AIcs.CVcs.CYcs.HCcs.CRcs.SDeess.ASphysics.comp-ph
On Valency
published · living versionsW_8zqhdax4·v1 · currentpublished
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
with Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud +9
1 version
Preprints & journals
29 papers in the corpus · 2019–2026Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet2605.29358v1 · Adly Templeton, Tom Conerly, Jonathan Marcus et al.2026 · 0 citationsarXiv
How AI Impacts Skill Formation2601.20245v2 · Judy Hanwen Shen, Alex Tamkin2026 · 4 citationsarXiv
Clio: Privacy-Preserving Insights into Real-World AI Use2412.13678v1 · Alex Tamkin, Miles McCain, Kunal Handa et al.2024 · 8 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Collective Constitutional AI: Aligning a Language Model with Public Input2406.07814v1 · Saffron Huang, Divya Siddarth, Liane Lovitt et al.2024 · 75 citationsProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1395-1417
Towards Measuring the Representation of Subjective Global Opinions in Language Models2306.16388v2 · Esin Durmus, Karina Nguyen, Thomas I. Liao et al.2023 · 43 citationsarXiv
Bayesian Preference Elicitation with Language Models2403.05534v1 · Kunal Handa, Yarin Gal, Ellie Pavlick et al.2024 · 0 citationsarXiv
Oolong: Investigating What Makes Transfer Learning Hard with Controlled Studies2202.12312v2 · Zhengxuan Wu, Alex Tamkin, Isabel Papadimitriou2022 · 10 citationsarXiv
Evaluating and Mitigating Discrimination in Language Model Decisions2312.03689v1 · Alex Tamkin, Amanda Askell, Liane Lovitt et al.2023 · 10 citationsarXiv
Social Contract AI: Aligning AI Assistants with Implicit Group Norms2310.17769v2 · Jan-Philipp Franken, Sam Kwok, Peixuan Ye et al.2023 · 1 citationarXiv
Turbulence in Focus: Benchmarking Scaling Behavior of 3D Volumetric Super-Resolution with BLASTNet 2.0 Data2309.13457v3 · Wai Tong Chung, Bassem Akoush, Pushan Sharma et al.2023 · 11 citationsarXiv
Codebook Features: Sparse and Discrete Interpretability for Neural Networks2310.17230v1 · Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman2023 · 2 citationsarXiv
Eliciting Human Preferences with Language Models2310.11589v1 · Belinda Z. Li, Alex Tamkin, Noah Goodman et al.2023 · 9 citationsarXiv
Studying Large Language Model Generalization with Influence Functions2308.03296v1 · Roger Grosse, Juhan Bae, Cem Anil et al.2023 · 26 citationsarXiv
BenchMD: A Benchmark for Unified Learning on Medical Images and Sensors2304.08486v2 · Kathryn Wantlin, Chenwei Wu, Shih-Cheng Huang et al.2023 · 10 citationsarXiv
Multispectral Contrastive Learning with Viewmaker Networks2302.05757v3 · Jasmine Bayrooti, Noah Goodman, Alex Tamkin2023 · 3 citationsarXiv
Operationalising the Definition of General Purpose AI Systems: Assessing Four Approaches2306.02889v1 · Risto Uuk, Carlos Ignacio Gutierrez, Alex Tamkin2023 · 3 citationsarXiv
DABS: A Domain-Agnostic Benchmark for Self-Supervised Learning2111.12062v2 · Alex Tamkin, Vincent Liu, Rongfei Lu et al.2021 · 6 citationsarXiv
Task Ambiguity in Humans and Language Models2212.10711v1 · Alex Tamkin, Kunal Handa, Avash Shrestha et al.2022 · 8 citationsarXiv
Feature Dropout: Revisiting the Role of Augmentations in Contrastive Learning2212.08378v1 · Alex Tamkin, Margalit Glasgow, Xiluo He et al.2022 · 1 citationarXiv
On the Opportunities and Risks of Foundation Models2108.07258v3 · Rishi Bommasani, Drew A. Hudson, Ehsan Adeli et al.2021 · 2,296 citationsarXiv
Active Learning Helps Pretrained Models Learn the Intended Task2204.08491v1 · Alex Tamkin, Dat Nguyen, Salil Deshpande et al.2022 · 14 citationsarXiv
Tradeoffs Between Contrastive and Supervised Learning: An Empirical Study2112.05340v1 · Ananya Karthik, Mike Wu, Noah Goodman et al.2021 · 0 citationsarXiv
C5T5: Controllable Generation of Organic Molecules with Transformers2108.10307v1 · Daniel Rothchild, Alex Tamkin, Julie Yu et al.2021 · 17 citationsarXiv
Viewmaker Networks: Learning Views for Unsupervised Representation Learning2010.07432v2 · Alex Tamkin, Mike Wu, Noah Goodman2020 · 29 citationsarXiv
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models2102.02503v1 · Alex Tamkin, Miles Brundage, Jack Clark et al.2021 · 129 citationsarXiv
Investigating Transferability in Pretrained Language Models2004.14975v2 · Alex Tamkin, Trisha Singh, Davide Giovanardi et al.2020 · 48 citationsarXiv
Language Through a Prism: A Spectral Approach for Multiscale Language Representations2011.04823v1 · Alex Tamkin, Dan Jurafsky, Noah Goodman2020 · 26 citationsarXiv
Being Optimistic to Be Conservative: Quickly Learning a CVaR Policy1911.01546v2 · Ramtin Keramati, Christoph Dann, Alex Tamkin et al.2019 · 38 citationsThirty-Fourth AAAI Conference on Artificial Intelligence (AAAI 2020)
Career total: 47 works. 29 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.