JL
Jan Leike
cs.AIcs.LGcs.CLstat.MLcs.LOcs.CRcs.NEcs.CYmath.STstat.TH
On Valency
published · living versionsW_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato +10
1 version
Preprints & journals
50 papers in the corpus · 2013–2026Unsupervised Elicitation of Language Models2506.10139v2 · Jiaxin Wen, Zachary Ankner, Arushi Somani et al.2025 · 0 citationsarXiv
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks2601.04603v1 · Hoagy Cunningham, Jerry Wei, Zihan Wang et al.2026 · 0 citationsarXiv
Excess Description Length of Learning Generalizable Predictors2601.04728v1 · Elizabeth Donoway, Hailey Joren, Fabien Roger et al.2026 · 0 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Forecasting Rare Language Model Behaviors2502.16797v1 · Erik Jones, Meg Tong, Jesse Mu et al.2025 · 0 citationsarXiv
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming2501.18837v1 · Mrinank Sharma, Meg Tong, Jesse Mu et al.2025 · 7 citationsarXiv
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Prover-Verifier Games improve legibility of LLM outputs2407.13692v2 · Jan Hendrik Kirchner, Yining Chen, Harri Edwards et al.2024 · 2 citationsarXiv
LLM Critics Help Catch LLM Bugs2407.00215v1 · Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe et al.2024 · 8 citationsarXiv
Scaling and evaluating sparse autoencoders2406.04093v1 · Leo Gao, Tom Dupr'e la Tour, Henk Tillman et al.2024 · 11 citationsarXivon Valency
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision2312.09390v1 · Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner et al.2023 · 25 citationsarXiv
Let's Verify Step by Step2305.20050v1 · Hunter Lightman, Vineet Kosaraju, Yura Burda et al.2023 · 31 citationsarXiv
Deep reinforcement learning from human preferences1706.03741v4 · Paul Christiano, Jan Leike, Tom B. Brown et al.2017 · 500 citationsarXiv
Self-critiquing models for assisting human evaluators2206.05802v2 · William Saunders, Catherine Yeh, Jeff Wu et al.2022 · 46 citationsarXiv
Training language models to follow instructions with human feedback2203.02155v1 · Long Ouyang, Jeff Wu, Xu Jiang et al.2022 · 5,565 citationsarXiv
Safe Deep RL in 3D Environments using Human Feedback2201.08102v2 · Matthew Rahtz, Vikrant Varma, Ramana Kumar et al.2022 · 2 citationsarXiv
Recursively Summarizing Books with Human Feedback2109.10862v2 · Jeff Wu, Long Ouyang, Daniel M. Ziegler et al.2021 · 68 citationsarXiv
Evaluating Large Language Models Trained on Code2107.03374v2 · Mark Chen, Jerry Tworek, Heewoo Jun et al.2021 · 1,468 citationsarXivon Valency
Learning Human Objectives by Evaluating Hypothetical Behavior1912.05652v2 · Siddharth Reddy, Anca D. Dragan, Sergey Levine et al.2019 · 9 citationsarXiv
Quantifying Differences in Reward Functions2006.13900v3 · Adam Gleave, Michael Dennis, Shane Legg et al.2020 · 15 citationsarXiv
Active Reinforcement Learning: Observing Rewards at a Cost2011.06709v2 · David Krueger, Jan Leike, Owain Evans et al.2020 · 12 citationsarXiv
Hidden Incentives for Auto-Induced Distributional Shift2009.09153v1 · David Krueger, Tegan Maharaj, Jan Leike2020 · 6 citationsarXiv
Pitfalls of learning a reward function online2004.13654v1 · Stuart Armstrong, Jan Leike, Laurent Orseau et al.2020 · 7 citationsarXiv
Learning to Understand Goal Specifications by Modelling Reward1806.01946v4 · Dzmitry Bahdanau, Felix Hill, Jan Leike et al.2018 · 62 citationsarXiv
Ranking Templates for Linear Loops1503.00193v2 · Jan Leike, Matthias Heizmann2015 · 39 citationsLogical Methods in Computer Science, Volume 11, Issue 1 (March 31, 2015) lmcs:797
Scaling shared model governance via model splitting1812.05979v1 · Miljan Martic, Jan Leike, Andrew Trask et al.2018 · 1 citationarXiv
Scalable agent alignment via reward modeling: a research direction1811.07871v1 · Jan Leike, David Krueger, Tom Everitt et al.2018 · 120 citationsarXiv
Reward learning from human preferences and demonstrations in Atari1811.06521v1 · Borja Ibarz, Jan Leike, Tobias Pohlen et al.2018 · 39 citationsarXiv
AI Safety Gridworlds1711.09883v2 · Jan Leike, Miljan Martic, Victoria Krakovna et al.2017 · 111 citationsarXiv
Universal Reinforcement Learning Algorithms: Survey and Experiments1705.10557v1 · John Aslanides, Jan Leike, Marcus Hutter2017 · 15 citationsarXiv
Generalised Discount Functions applied to a Monte-Carlo AImu Implementation1703.01358v1 · Sean Lamont, John Aslanides, Jan Leike et al.2017 · 0 citationsarXiv
Nonparametric General Reinforcement Learning1611.08944v1 · Jan Leike2016 · 6 citationsarXiv
Exploration Potential1609.04994v3 · Jan Leike2016 · 2 citationsarXiv
A Formal Solution to the Grain of Truth Problem1609.05058v1 · Jan Leike, Jessica Taylor, Benya Fallenstein2016 · 5 citationsarXiv
Geometric Nontermination Arguments1609.05207v1 · Jan Leike, Matthias Heizmann2016 · 27 citationsarXiv
Thompson Sampling is Asymptotically Optimal in General Environments1602.07905v2 · Jan Leike, Tor Lattimore, Laurent Orseau et al.2016 · 23 citationsarXiv
Loss Bounds and Time Complexity for Speed Priors1604.03343v1 · Daniel Filan, Marcus Hutter, Jan Leike2016 · 1 citationarXiv
Solomonoff Induction Violates Nicod's Criterion1507.04121v1 · Jan Leike, Marcus Hutter2015 · 2 citationsarXiv
On the Computability of Solomonoff Induction and Knowledge-Seeking1507.04124v1 · Jan Leike, Marcus Hutter2015 · 8 citationsarXiv
Bad Universal Priors and Notions of Optimality1510.04931v1 · Jan Leike, Marcus Hutter2015 · 56 citationsarXiv
On the Computability of AIXI1510.05572v1 · Jan Leike, Marcus Hutter2015 · 2 citationsarXiv
Sequential Extensions of Causal and Evidential Decision Theory1506.07359v1 · Tom Everitt, Jan Leike, Marcus Hutter2015 · 9 citationsarXiv
A Definition of Happiness for Reinforcement Learning Agents1505.04497v1 · Mayank Daswani, Jan Leike2015 · 9 citationsarXiv
Indefinitely Oscillating Martingales1408.3169v1 · Jan Leike, Marcus Hutter2014 · 1 citationarXiv
Geometric Series as Nontermination Arguments for Linear Lasso Programs1405.4413v1 · Jan Leike, Matthias Heizmann2014 · 4 citationsarXiv
Ranking Templates for Linear Loops1401.5338v1 · Jan Leike, Matthias Heizmann2014 · 43 citationsarXiv
Linear Ranking for Linear Lasso Programs1401.5347v1 · Matthias Heizmann, Jochen Hoenicke, Jan Leike et al.2014 · 50 citationsarXiv
Ranking Function Synthesis for Linear Lasso Programs1401.5351v1 · Jan Leike2014 · 5 citationsarXiv
Synthesis for Polynomial Lasso Programs1311.4046v1 · Jan Leike, Ashish Tiwari2013 · 5 citationsarXiv
Career total: 70 works. 50 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.