PA

Pieter Abbeel

cs.LGcs.AIcs.ROcs.CVstat.MLcs.NEcs.CLcs.SYeess.SYcs.GR

On Valency

published · living versions
W_7f666ryv·v1 · currentpublished
Benchmarking Deep Reinforcement Learning for Continuous Control
with Yan Duan, Xi Chen, Rein Houthooft, John Schulman
1 version

Preprints & journals

390 papers in the corpus · 2012–2026
Evaluating Protein Transfer Learning with TAPE.33390682 · Rao, Roshan, Bhattacharya, Nicholas, Thomas, Neil et al.2026 · 752 citationsAdvances in neural information processing systems. 2019;32:9689-9701
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes2606.30645v1 · Yen-Jen Wang, Jiaman Li, Sirui Chen et al.2026 · 0 citationsarXiv
Scalable Behavior Cloning with Open Data, Training, and Evaluation2606.27375v1 · Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh et al.2026 · 0 citationsarXiv
Do as I Do: Dexterous Manipulation Data from Everyday Human Videos2606.19333v1 · Bhawna Paliwal, Haritheja Etukuru, William Liang et al.2026 · 0 citationsarXiv
SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation2606.10305v1 · Qianzhong Chen, Hau Zheng, Justin Yu et al.2026 · 0 citationsarXiv
High-Dimensional Continuous Control Using Generalized Advantage Estimation1506.02438v6 · John Schulman, Philipp Moritz, Sergey Levine et al.2015 · 1,817 citationsarXiv
Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model1610.03518v1 · Paul Christiano, Zain Shah, Igor Mordatch et al.2016 · 161 citationsarXiv
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization2511.22445v2 · Yikai Tang, Haoran Geng, Jindou Jia et al.2025 · 0 citationsarXiv
Reward-Conditioned Reinforcement Learning2603.05066v3 · Michal Nauman, Marek Cygan, Pieter Abbeel2026 · 0 citationsarXiv
When Does Non-Uniform Replay Matter in Reinforcement Learning?2605.10236v3 · Michal Korniak, Miko\laj Czarnecki, Yarden As et al.2026 · 0 citationsarXiv
Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching2602.15827v2 · Zhen Wu, Xiaoyu Huang, Lujie Yang et al.2026 · 0 citationsarXiv
World Model for Robot Learning: A Comprehensive Survey2605.00080v1 · Bohan Hou, Gen Li, Jindou Jia et al.2026 · 0 citationsarXiv
Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own2310.02635v5 · Weirui Ye, Yunsheng Zhang, Haoyang Weng et al.2023 · 0 citationsarXiv
Rodrigues Network for Learning Robot Actions2506.02618v2 · Jialiang Zhang, Haoran Geng, Yang You et al.2025 · 0 citationsarXiv
Relative Entropy Pathwise Policy Optimization2507.11019v4 · Claas Voelcker, Axel Brunnbauer, Marcel Hussing et al.2025 · 0 citationsarXiv
OSGym: Scalable OS Infra for Computer Use Agents2511.11672v5 · Zengyi Qin, Jinyuan Chen, Yunze Man et al.2025 · 0 citationsarXiv
Cross-Hand Latent Representation for Vision-Language-Action Models2603.10158v1 · Guangqi Jiang, Yutong Liang, Jianglong Ye et al.2026 · 0 citationsarXiv
Controllable All-Atom Protein Generation with Latent Diffusion10.1101/2024.12.02.626353v3 · Lu, A. X., Yan, W., Robinson, S. A. et al.2024 · 5 citationsbioRxiv
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction2503.03734v4 · Huang Huang, Fangchen Liu, Letian Fu et al.2025 · 0 citationsarXiv
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes2512.01996v1 · Younggyo Seo, Carmelo Sferrazza, Juyue Chen et al.2025 · 0 citationsarXiv
Learning to Design Soft Hands using Reward Models2510.17086v1 · Xueqian Bai, Nicklas Hansen, Adabhav Singh et al.2025 · 0 citationsarXiv
GaussGym: An open-source real-to-sim framework for learning locomotion from pixels2510.15352v1 · Alejandro Escontrela, Justin Kerr, Arthur Allshire et al.2025 · 0 citationsarXiv
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy2505.11032v3 · Yuran Wang, Ruihai Wu, Yue Chen et al.2025 · 1 citationarXiv
Residual Off-Policy RL for Finetuning Behavior Cloning Policies2509.19301v2 · Lars Ankile, Zhenyu Jiang, Rocky Duan et al.2025 · 0 citationsarXiv
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio2507.02864v2 · Renhao Wang, Haoran Geng, Tingle Li et al.2025 · 0 citationsarXiv
End-to-end RL Improves Dexterous Grasping Policies2509.16434v1 · Ritvik Singh, Karl Van Wyk, Pieter Abbeel et al.2025 · 0 citationsarXiv
ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface2504.06156v2 · Fangchen Liu, Chuanyu Li, Yihua Qin et al.2025 · 0 citationsarXiv
Compute-Optimal Scaling for Value-Based Deep RL2508.14881v2 · Preston Fu, Oleh Rybkin, Zhiyuan Zhou et al.2025 · 1 citationarXiv
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization2412.12098v2 · Bhavya Sukhija, Stelian Coros, Andreas Krause et al.2024 · 1 citationarXiv
Value-Based Deep RL Scales Predictably2502.04327v2 · Oleh Rybkin, Michal Nauman, Preston Fu et al.2025 · 0 citationsarXiv
From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control2405.04798v3 · Yide Shentu, Philipp Wu, Aravind Rajeswaran et al.2024 · 11 citationsarXiv
Tokenized and continuous embedding compressions of protein sequence and structure.40575127 · Lu, Amy X, Yan, Wilson, Yang, Kevin K et al.2025 · 13 citationsPatterns (New York, N.Y.). 2025;6(6):101289
One Step Diffusion via Shortcut Models2410.12557v3 · Kevin Frans, Danijar Hafner, Sergey Levine et al.2024 · 2 citationsarXiv
SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending2506.09366v1 · Yuxuan Kuang, Haoran Geng, Amine Elhafsi et al.2025 · 0 citationsarXiv
Chip Placement with Diffusion Models2407.12282v3 · Vint Lee, Minh Nguyen, Leena Elzeiny et al.2024 · 0 citationsarXiv
EgoZero: Robot Learning from Smart Glasses2505.20290v2 · Vincent Liu, Ademi Adeniji, Haotian Zhan et al.2025 · 0 citationsarXiv
Object-centric 3D Motion Field for Robot Learning from Human Videos2506.04227v1 · Zhao-Heng Yin, Sherry Yang, Pieter Abbeel2025 · 0 citationsarXiv
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control2505.22642v3 · Younggyo Seo, Carmelo Sferrazza, Haoran Geng et al.2025 · 0 citationsarXiv
Feel the Force: Contact-Driven Learning from Humans2506.01944v1 · Ademi Adeniji, Zhuoran Chen, Vincent Liu et al.2025 · 0 citationsarXiv
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners2505.23150v1 · Michal Nauman, Marek Cygan, Carmelo Sferrazza et al.2025 · 0 citationsarXiv
Diffusion Guidance Is a Controllable Policy Improvement Operator2505.23458v1 · Kevin Frans, Seohong Park, Pieter Abbeel et al.2025 · 0 citationsarXiv
Prioritized Generative Replay2410.18082v2 · Renhao Wang, Kevin Frans, Pieter Abbeel et al.2024 · 0 citationsarXiv
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction2411.14762v4 · Huiwon Jang, Sihyun Yu, Jinwoo Shin et al.2024 · 0 citationsarXiv
RoboCopilot: Human-in-the-loop Interactive Imitation Learning for Robot Manipulation2503.07771v1 · Philipp Wu, Yide Shentu, Qiayuan Liao et al.2025 · 1 citationarXiv
Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm2503.07541v1 · Zhao-Heng Yin, Changhao Wang, Luis Pineda et al.2025 · 6 citationsarXiv
Career total: 658 works. 390 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.