SL

Sergey Levine

cs.LGcs.ROcs.AIstat.MLcs.CVcs.CLcs.NEcs.SYeess.SYcs.HC

On Valency

published · living versions
W_9kmp5csx·v1 · currentpublished
Trust Region Policy Optimization
with John Schulman, Philipp Moritz, Michael I. Jordan, Pieter Abbeel
1 version

Preprints & journals

529 papers in the corpus · 2012–2026
Leveraging Discrete Function Decomposability for Scientific Design2511.03032v3 · James C. Bowden, Sergey Levine, Jennifer Listgarten2025 · 0 citationsarXiv
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models2508.13446v2 · Catherine Glossop, William Chen, Arjun Bhorkar et al.2025 · 0 citationsarXiv
High-Dimensional Continuous Control Using Generalized Advantage Estimation1506.02438v6 · John Schulman, Philipp Moritz, Sergey Levine et al.2015 · 1,817 citationsarXiv
Continuous Deep Q-Learning with Model-based Acceleration1603.00748v1 · Shixiang Gu, Timothy Lillicrap, Ilya Sutskever et al.2016 · 320 citationsarXiv
Learning Dexterous Manipulation Policies from Experience and Imitation1611.05095v1 · Vikash Kumar, Abhishek Gupta, Emanuel Todorov et al.2016 · 40 citationsarXiv
Goal-Driven Dynamics Learning via Bayesian Optimization1703.09260v2 · Somil Bansal, Roberto Calandra, Ted Xiao et al.2017 · 89 citationsarXiv
MBMF: Model-Based Priors for Model-Free Reinforcement Learning1709.03153v2 · Somil Bansal, Roberto Calandra, Kurtland Chua et al.2017 · 20 citationsarXiv
Intention-Conditioned Flow Occupancy Models2506.08902v4 · Chongyi Zheng, Seohong Park, Sergey Levine et al.2025 · 0 citationsarXiv
Reinforcement Learning with Action Chunking2507.07969v4 · Qiyang Li, Zhiyuan Zhou, Sergey Levine2025 · 1 citationarXiv
Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning2511.07730v3 · Bill Chunyuan Zheng, Vivek Myers, Benjamin Eysenbach et al.2025 · 0 citationsICLR (2026)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space2512.04601v2 · Joey Hong, Kang Liu, Zhan Ling et al.2025 · 1 citationarXiv
$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control2410.24164v4 · Kevin Black, Noah Brown, Danny Driess et al.2024 · 11 citationsarXiv
RoboReward: General-Purpose Vision-Language Reward Models for Robotics2601.00675v2 · Tony Lee, Andrew Wagenmaker, Karl Pertsch et al.2026 · 0 citationsarXiv
PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies2512.16881v2 · Arhan Jain, Mingtong Zhang, Kanav Arora et al.2025 · 2 citationsarXiv
Emergence of Human to Robot Transfer in Vision-Language-Action Models2512.22414v1 · Simar Kareer, Karl Pertsch, James Darpinian et al.2025 · 1 citationarXiv
Zero-Overhead Introspection for Adaptive Test-Time Compute2512.01457v4 · Rohin Manvi, Joey Hong, Tim Seyde et al.2025 · 0 citationsarXiv
Decoupled Q-Chunking2512.10926v2 · Qiyang Li, Seohong Park, Sergey Levine2025 · 0 citationsarXiv
Training-Time Action Conditioning for Efficient Real-Time Chunking2512.05964v2 · Kevin Black, Allen Z. Ren, Michael Equi et al.2025 · 0 citationsarXiv
Real-Time Execution of Action Chunking Flow Policies2506.07339v2 · Kevin Black, Manuel Y. Galliker, Sergey Levine2025 · 7 citationsarXiv
Learning to Drive Anywhere with Model-Based Reannotation2505.05592v3 · Noriaki Hirose, Lydia Ignatova, Kyle Stachowicz et al.2025 · 0 citationsIEEE Robotics and Automation Letters 2025
$\pi^{*}_{0.6}$: a VLA That Learns From Experience2511.14759v2 · Physical Intelligence, Ali Amin, Raichelle Aniceto et al.2025 · 0 citationsarXiv
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning2511.00222v1 · Marwa Abdulhai, Ryan Cheng, Donovan Clay et al.2025 · 1 citationarXiv
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks2511.01758v1 · Mian Wu, Gavin Zhang, Sewon Min et al.2025 · 0 citationsarXiv
Horizon Reduction Makes RL Scalable2506.04168v3 · Seohong Park, Kevin Frans, Deepinder Mann et al.2025 · 0 citationsarXiv
Learning Affordances at Inference-Time for Vision-Language-Action Models2510.19752v1 · Ameesh Shah, William Chen, Adwait Godbole et al.2025 · 0 citationsarXiv
Training LLM Agents to Empower Humans2510.13709v2 · Evan Ellis, Vivek Myers, Jens Tuyls et al.2025 · 0 citationsarXiv
Compute-Optimal Scaling for Value-Based Deep RL2508.14881v2 · Preston Fu, Oleh Rybkin, Zhiyuan Zhou et al.2025 · 1 citationarXiv
Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.40834062 · Luo, Jianlan, Xu, Charles, Wu, Jeffrey et al.2025 · 67 citationsScience robotics. 2025;10(105):eads5033
Value-Based Deep RL Scales Predictably2502.04327v2 · Oleh Rybkin, Michal Nauman, Preston Fu et al.2025 · 0 citationsarXiv
Adapt On-the-Go: Behavior Modulation for Single-Life Robot Deployment2311.01059v3 · Annie S. Chen, Govind Chada, Laura Smith et al.2023 · 2 citationsarXiv
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models2502.19417v2 · Lucy Xiaoyang Shi, Brian Ichter, Michael Equi et al.2025 · 0 citationsarXiv
Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration2410.18076v4 · Max Wilcoxson, Qiyang Li, Kevin Frans et al.2024 · 0 citationsarXiv
Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data2412.07762v3 · Zhiyuan Zhou, Andy Peng, Qiyang Li et al.2024 · 1 citationInternational Conference on Learning Representation (ICLR) 2025
Steering Your Diffusion Policy with Latent Space Reinforcement Learning2506.15799v2 · Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang et al.2025 · 0 citationsarXiv
One Step Diffusion via Shortcut Models2410.12557v3 · Kevin Frans, Danijar Hafner, Sergey Levine et al.2024 · 2 citationsarXiv
Visual Pre-Training on Unlabeled Images using Reinforcement Learning2506.11967v1 · Dibya Ghosh, Sergey Levine2025 · 0 citationsarXiv
Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data2306.03346v3 · Chongyi Zheng, Benjamin Eysenbach, Homer Walke et al.2023 · 2 citationsarXiv
Dynamic Search for Inference-Time Alignment in Diffusion Models2503.02039v2 · Xiner Li, Masatoshi Uehara, Xingyu Su et al.2025 · 0 citationsarXiv
Self-Challenging Language Model Agents2506.01716v1 · Yifei Zhou, Sergey Levine, Jason Weston et al.2025 · 1 citationarXiv
Diffusion Guidance Is a Controllable Policy Improvement Operator2505.23458v1 · Kevin Frans, Seohong Park, Pieter Abbeel et al.2025 · 0 citationsarXiv
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better2505.23705v1 · Danny Driess, Jost Tobias Springenberg, Brian Ichter et al.2025 · 4 citationsarXiv
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training2501.17161v2 · Tianzhe Chu, Yuexiang Zhai, Jihan Yang et al.2025 · 6 citationsarXiv
Flow Q-Learning2502.02538v2 · Seohong Park, Qiyang Li, Sergey Levine2025 · 0 citationsarXiv
Inference via Interpolation: Contrastive Representations Provably Enable Planning and Inference2403.04082v4 · Benjamin Eysenbach, Vivek Myers, Ruslan Salakhutdinov et al.2024 · 1 citationNeural information processing systems (2024)
Open X-Embodiment: Robotic Learning Datasets and RT-X Models2310.08864v9 · Abby O'Neill, Abdul Rehman, Abhinav Gupta et al.2023 · 104 citationsarXiv
Prioritized Generative Replay2410.18082v2 · Renhao Wang, Kevin Frans, Pieter Abbeel et al.2024 · 0 citationsarXiv
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset2403.12945v2 · Alexander Khazatsky, Karl Pertsch, Suraj Nair et al.2024 · 143 citationsarXiv
$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization2504.16054v1 · Physical Intelligence, Kevin Black, Noah Brown et al.2025 · 3 citationsarXiv
Career total: 751 works. 529 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.