JK

Jared Kaplan

hep-thcs.LGcs.CLcs.AIhep-phcond-mat.str-elcs.CYstat.MLcs.CRhep-ex

On Valency

published · living versions
W_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato +10
1 version

Preprints & journals

68 papers in the corpus · 2007–2025
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Forecasting Rare Language Model Behaviors2502.16797v1 · Erik Jones, Meg Tong, Jesse Mu et al.2025 · 0 citationsarXiv
Alignment faking in large language models2412.14093v2 · Ryan Greenblatt, Carson Denison, Benjamin Wright et al.2024 · 26 citationsarXivon Valency
Clio: Privacy-Preserving Insights into Real-World AI Use2412.13678v1 · Alex Tamkin, Miles McCain, Kunal Handa et al.2024 · 8 citationsarXiv
Sabotage Evaluations for Frontier Models2410.21514v1 · Joe Benton, Misha Wagner, Eric Christiansen et al.2024 · 2 citationsarXiv
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models2406.10162v3 · Carson Denison, Monte MacDiarmid, Fazl Barez et al.2024 · 9 citationsarXivon Valency
Explaining Neural Scaling Laws2102.06701v2 · Yasaman Bahri, Ethan Dyer, Jared Kaplan et al.2021 · 175 citationsPNAS 121 (27) e2311878121 (2024)
Towards Measuring the Representation of Subjective Global Opinions in Language Models2306.16388v2 · Esin Durmus, Karina Nguyen, Thomas I. Liao et al.2023 · 43 citationsarXiv
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2401.05566v3 · Evan Hubinger, Carson Denison, Jesse Mu et al.2024 · 39 citationsarXivon Valency
Evaluating and Mitigating Discrimination in Language Model Decisions2312.03689v1 · Alex Tamkin, Amanda Askell, Liane Lovitt et al.2023 · 10 citationsarXiv
Specific versus General Principles for Constitutional AI2310.13798v1 · Sandipan Kundu, Yuntao Bai, Saurav Kadavath et al.2023 · 7 citationsarXiv
Studying Large Language Model Generalization with Influence Functions2308.03296v1 · Roger Grosse, Juhan Bae, Cem Anil et al.2023 · 26 citationsarXiv
Measuring Faithfulness in Chain-of-Thought Reasoning2307.13702v1 · Tamera Lanham, Anna Chen, Ansh Radhakrishnan et al.2023 · 32 citationsarXiv
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning2307.11768v2 · Ansh Radhakrishnan, Karina Nguyen, Anna Chen et al.2023 · 8 citationsarXiv
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models2206.04615v3 · Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al.2022 · 565 citationsTransactions on Machine Learning Research, May/2022, https://openreview.net/forum?id=uyTL5Bvosj
The Capacity for Moral Self-Correction in Large Language Models2302.07459v2 · Deep Ganguli, Amanda Askell, Nicholas Schiefer et al.2023 · 53 citationsarXiv
Discovering Language Model Behaviors with Model-Written Evaluations2212.09251v1 · Ethan Perez, Sam Ringer, Kamil\.e Lukosi\=ut\.e et al.2022 · 224 citationsarXivon Valency
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
The Bulk-to-Boundary Propagator in Black Hole Microstate Backgrounds1810.02436v2 · Hongbin Chen, A. Liam Fitzpatrick, Jared Kaplan et al.2018 · 15 citationsarXiv
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
In-context Learning and Induction Heads2209.11895v1 · Catherine Olsson, Nelson Elhage, Neel Nanda et al.2022 · 87 citationsarXivon Valency
Toy Models of Superposition2209.10652v1 · Nelson Elhage, Tristan Hume, Catherine Olsson et al.2022 · 49 citationsarXiv
Scaling Laws and Interpretability of Learning from Repeated Data2205.10487v1 · Danny Hernandez, Tom Brown, Tom Conerly et al.2022 · 22 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
A General Language Assistant as a Laboratory for Alignment2112.00861v3 · Amanda Askell, Yuntao Bai, Anna Chen et al.2021 · 27 citationsarXiv
Causality Constraints in Large $N$ QCD Coupled to Gravity2009.08460v2 · Jared Kaplan, Sandipan Kundu2020 · 11 citationsPhys. Rev. D 104, 061901 (2021)
Evaluating Large Language Models Trained on Code2107.03374v2 · Mark Chen, Jerry Tworek, Heewoo Jun et al.2021 · 1,468 citationsarXivon Valency
Scaling Laws for Transfer2102.01293v1 · Danny Hernandez, Jared Kaplan, Tom Henighan et al.2021 · 26 citationsarXiv
Scaling Laws for Autoregressive Generative Modeling2010.14701v2 · Tom Henighan, Jared Kaplan, Mor Katz et al.2020 · 150 citationsarXiv
Language Models are Few-Shot Learners2005.14165v4 · Tom B. Brown, Benjamin Mann, Nick Ryder et al.2020 · 2,966 citationsarXiv
A Neural Scaling Law from the Dimension of the Data Manifold2004.10802v1 · Utkarsh Sharma, Jared Kaplan2020 · 20 citationsarXiv
An Empirical Model of Large-Batch Training1812.06162v1 · Sam McCandlish, Jared Kaplan, Dario Amodei et al.2018 · 129 citationsarXiv
A Numerical Approach to Virasoro Blocks and the Information Paradox1703.09727v1 · Hongbin Chen, Charles Hussong, Jared Kaplan et al.2017 · 82 citationsarXiv
On Information Loss in AdS$_3$/CFT$_2$1603.08925v3 · A. Liam Fitzpatrick, Jared Kaplan, Daliang Li et al.2016 · 27 citationsarXiv
Hawking from Catalan1510.00014v1 · A. Liam Fitzpatrick, Jared Kaplan, Matthew T. Walters et al.2015 · 84 citationsJHEP 05 (2016) 069
Conformal Blocks Beyond the Semi-Classical Limit1512.03052v3 · A. Liam Fitzpatrick, Jared Kaplan2015 · 100 citationsarXiv
A Quantum Correction To Chaos1601.06164v1 · A. Liam Fitzpatrick, Jared Kaplan2016 · 82 citationsarXiv
Virasoro Conformal Blocks and Thermality from Classical Background Fields1501.05315v3 · A. Liam Fitzpatrick, Jared Kaplan, Matthew T. Walters2015 · 329 citationsJHEP 11 (2015) 200
Eikonalization of Conformal Blocks1504.01737v3 · A. Liam Fitzpatrick, Jared Kaplan, Matthew T. Walters et al.2015 · 82 citationsJHEP 09 (2015) 019
AdS Field Theory from Conformal Field Theory1208.0337v1 · A. Liam Fitzpatrick, Jared Kaplan2012 · 119 citationsarXiv
Analyticity and the Holographic S-Matrix1111.6972v3 · A. Liam Fitzpatrick, Jared Kaplan2011 · 212 citationsarXiv
Unraveling L_{n,k}: Grassmannian Kinematics0912.0957v1 · Jared Kaplan2009 · 46 citationsarXiv
The S-Matrix in Twistor Space0903.2110v2 · Nima Arkani-Hamed, Freddy Cachazo, Clifford Cheung et al.2009 · 178 citationsarXiv
A Natural Language for AdS/CFT Correlators1107.1499v2 · A. Liam Fitzpatrick, Jared Kaplan, Joao Penedones et al.2011 · 273 citationsJHEP Vol. 2011, No. 11, 95
LHC Predictions from a Tevatron Anomaly in the Top Quark Forward-Backward Asymmetry1101.5203v1 · Yang Bai, JoAnne L. Hewett, Jared Kaplan et al.2011 · 120 citationsJHEP 1103:003,2011
Career total: 87 works. 68 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.