JS

John Schulman

cs.LGcs.AIstat.MLcs.CLcs.CVcs.ROcs.CRcs.CYcs.NEcs.SD

On Valency

published · living versions
W_sgh3nvvu·v1 · currentpublished
Reasoning Models Don't Always Say What They Think
with Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato +10
1 version

Preprints & journals

31 papers in the corpus · 2015–2025
High-Dimensional Continuous Control Using Generalized Advantage Estimation1506.02438v6 · John Schulman, Philipp Moritz, Sergey Levine et al.2015 · 1,817 citationsarXiv
Detecting Adversarial Fine-tuning with Auditing Agents2510.16255v1 · Sarah Egler, John Schulman, Nicholas Carlini2025 · 0 citationsarXiv
Reasoning Models Don't Always Say What They Think2505.05410v1 · Yanda Chen, Joe Benton, Ansh Radhakrishnan et al.2025 · 9 citationsarXivon Valency
Measuring short-form factuality in large language models2411.04368v1 · Jason Wei, Nguyen Karina, Hyung Won Chung et al.2024 · 13 citationsarXiv
Rule Based Rewards for Language Model Safety2411.01111v1 · Tong Mu, Alec Helyar, Johannes Heidecke et al.2024 · 10 citationsarXiv
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Unsolved Problems in ML Safety2109.13916v5 · Dan Hendrycks, Nicholas Carlini, John Schulman et al.2021 · 85 citationsarXiv
Training language models to follow instructions with human feedback2203.02155v1 · Long Ouyang, Jeff Wu, Xu Jiang et al.2022 · 5,565 citationsarXiv
Scaling Laws for Autoregressive Generative Modeling2010.14701v2 · Tom Henighan, Jared Kaplan, Mor Katz et al.2020 · 150 citationsarXiv
Phasic Policy Gradient2009.04416v1 · Karl Cobbe, Jacob Hilton, Oleg Klimov et al.2020 · 51 citationsarXiv
Teacher-Student Curriculum Learning.31502993 · Matiisen, Tambet, Oliver, Avital, Cohen, Taco et al.2020 · 324 citationsIEEE transactions on neural networks and learning systems. 2020;31(9):3732-3740
Leveraging Procedural Generation to Benchmark Reinforcement Learning1912.01588v2 · Karl Cobbe, Christopher Hesse, Jacob Hilton et al.2019 · 204 citationsarXiv
Quantifying Generalization in Reinforcement Learning1812.02341v3 · Karl Cobbe, Oleg Klimov, Chris Hesse et al.2018 · 189 citationsarXiv
Semi-Supervised Learning by Label Gradient Alignment1902.02336v1 · Jacob Jackson, John Schulman2019 · 15 citationsarXiv
On First-Order Meta-Learning Algorithms1803.02999v3 · Alex Nichol, Joshua Achiam, John Schulman2018 · 517 citationsarXiv
Equivalence Between Policy Gradients and Soft Q-Learning1704.06440v4 · John Schulman, Xi Chen, Pieter Abbeel2017 · 182 citationsarXiv
Model-Based Reinforcement Learning via Meta-Policy Optimization1809.05214v1 · Ignasi Clavera, Jonas Rothfuss, John Schulman et al.2018 · 110 citationsarXiv
Gotta Learn Fast: A New Benchmark for Generalization in RL1804.03720v2 · Alex Nichol, Vicki Pfau, Christopher Hesse et al.2018 · 80 citationsarXiv
#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning1611.04717v3 · Haoran Tang, Rein Houthooft, Davis Foote et al.2016 · 464 citationsarXiv
Teacher-Student Curriculum Learning1707.00183v2 · Tambet Matiisen, Avital Oliver, Taco Cohen et al.2017 · 324 citationsarXiv
UCB Exploration via Q-Ensembles1706.01502v3 · Richard Y. Chen, Szymon Sidor, Pieter Abbeel et al.2017 · 76 citationsarXiv
Meta Learning Shared Hierarchies1710.09767v1 · Kevin Frans, Jonathan Ho, Xi Chen et al.2017 · 111 citationsarXiv
Trust Region Policy Optimization1502.05477v5 · John Schulman, Sergey Levine, Philipp Moritz et al.2015 · 3,084 citationsarXivon Valency
Variational Lossy Autoencoder1611.02731v2 · Xi Chen, Diederik P. Kingma, Tim Salimans et al.2016 · 241 citationsarXiv
VIME: Variational Information Maximizing Exploration1605.09674v4 · Rein Houthooft, Xi Chen, Yan Duan et al.2016 · 521 citationsarXiv
RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning1611.02779v2 · Yan Duan, John Schulman, Xi Chen et al.2016 · 479 citationsarXiv
OpenAI Gym1606.01540v1 · Greg Brockman, Vicki Cheung, Ludwig Pettersson et al.2016 · 620 citationsarXiv
Benchmarking Deep Reinforcement Learning for Continuous Control1604.06778v3 · Yan Duan, Xi Chen, Rein Houthooft et al.2016 · 1,254 citationsarXivon Valency
Gradient Estimation Using Stochastic Computation Graphs1506.05254v3 · John Schulman, Nicolas Heess, Theophane Weber et al.2015 · 173 citationsarXiv
Career total: 54 works. 31 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.