PK
Philipp Krahenbuhl
cs.CVcs.LGcs.AIcs.ROcs.CLeess.IVstat.MLbioinformaticscs.GRcs.IT
On Valency
published · living versionsW_ber6cpqu·v1 · currentpublished
Context Encoders: Feature Learning by Inpainting
with Deepak Pathak, Jeff Donahue, Trevor Darrell, Alexei A. Efros
1 version
Preprints & journals
62 papers in the corpus · 2011–2026Latent Chain-of-Thought World Modeling for End-to-End Driving2512.10226v3 · Shuhan Tan, Kashyap Chitta, Yuxiao Chen et al.2025 · 0 citationsarXiv
Mask-Aware Policy Gradients for Diffusion Language Models2607.15200v1 · Haran Raajesh, Kulin Shah, Adam Klivans et al.2026 · 0 citationsarXiv
Do multimodal models imagine electric sheep?2605.09693v1 · Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes et al.2026 · 0 citationsarXiv
Entropy-Preserving Reinforcement Learning2603.11682v1 · Aleksei Petrenko, Ben Lipkin, Kevin Chen et al.2026 · 0 citationsProceedings of the International Conference on Learning Representations (ICLR), 2026
Compressed Map Priors for 3D Perception2601.00139v1 · Brady Zhou, Philipp Krahenbuhl2025 · 0 citationsarXiv
Spherical Leech Quantization for Visual Tokenization and Generation2512.14697v1 · Yue Zhao, Hanwen Jiang, Zhenlin Xu et al.2025 · 0 citationsarXiv
Triangle Multiplication Is All You Need For Biomolecular Structure Representations2510.18870v2 · Jeffrey Ouyang-Zhang, Pranav Murugan, Daniel J. Diaz et al.2025 · 0 citationsarXiv
Domain Adaptation Through Task Distillation2008.11911v1 · Brady Zhou, Nimit Kalra, Philipp Krahenbuhl2020 · 14 citationsarXiv
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation2506.17213v2 · Xiuyu Yang, Shuhan Tan, Philipp Krahenbuhl2025 · 1 citationarXiv
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding2504.13180v3 · Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi et al.2025 · 5 citationsarXiv
Interactive Post-Training for Vision-Language-Action Models2505.17016v1 · Shuhan Tan, Kairan Dou, Yue Zhao et al.2025 · 0 citationsarXiv
Cut Your Losses in Large-Vocabulary Language Models2411.09009v2 · Erik Wijmans, Brody Huval, Alexander Hertzberg et al.2024 · 2 citationsarXiv
Reinforcement Learning for Long-Horizon Interactive LLM Agents2502.01600v3 · Kevin Chen, Marco Cusumano-Towner, Brody Huval et al.2025 · 3 citationsarXiv
Distilling structural representations into protein sequence models10.1101/2024.11.08.622579v2 · Ouyang-Zhang, J., Gong, C., Zhao, Y. et al.2024 · 11 citationsbioRxiv
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation2502.05178v1 · Yue Zhao, Fuzhao Xue, Scott Reed et al.2025 · 0 citationsarXiv
Robust Autonomy Emerges from Self-Play2502.03349v1 · Marco Cusumano-Towner, David Hafner, Alex Hertzberg et al.2025 · 0 citationsarXiv
Promptable Closed-loop Traffic Simulation2409.05863v1 · Shuhan Tan, Boris Ivanovic, Yuxiao Chen et al.2024 · 0 citationsarXiv
Image and Video Tokenization with Binary Spherical Quantization2406.07548v1 · Yue Zhao, Yuanjun Xiong, Philipp Krahenbuhl2024 · 0 citationsarXiv
Language-Image Models with 3D Understanding2405.03685v1 · Jang Hyun Cho, Boris Ivanovic, Yulong Cao et al.2024 · 0 citationsarXiv
Distilling Vision-Language Models on Millions of Videos2401.06129v2 · Yue Zhao, Long Zhao, Xingyi Zhou et al.2024 · 13 citationsarXiv
Language-conditioned Detection Transformer2311.17902v1 · Jang Hyun Cho, Philipp Krahenbuhl2023 · 0 citationsarXiv
Predicting a Protein's Stability under a Million Mutations2310.12979v2 · Jeffrey Ouyang-Zhang, Daniel J. Diaz, Adam R. Klivans et al.2023 · 18 citationsarXiv
Training a Large Video Model on a Single Machine in a Day2309.16669v1 · Yue Zhao, Philipp Krahenbuhl2023 · 5 citationsarXiv
Long-tail Detection with Effective Class-Margins2301.09724v1 · Jang Hyun Cho, Philipp Krahenbuhl2023 · 17 citationsarXiv
NMS Strikes Back2212.06137v1 · Jeffrey Ouyang-Zhang, Jang Hyun Cho, Xingyi Zhou et al.2022 · 6 citationsarXiv
Learning Video Representations from Large Language Models2212.04501v1 · Yue Zhao, Ishan Misra, Philipp Krahenbuhl et al.2022 · 155 citationsarXiv
Real-time Online Video Detection with Temporal Smoothing Transformers2209.09236v1 · Yue Zhao, Philipp Krahenbuhl2022 · 73 citationsarXiv
Learning from All Vehicles2203.11934v3 · Dian Chen, Philipp Krahenbuhl2022 · 187 citationsarXiv
Detecting Twenty-thousand Classes using Image-level Supervision2201.02605v3 · Xingyi Zhou, Rohit Girdhar, Armand Joulin et al.2022 · 538 citationsarXiv
Cross-view Transformers for real-time Map-view Semantic Segmentation2205.02833v1 · Brady Zhou, Philipp Krahenbuhl2022 · 312 citationsarXiv
Simple multi-dataset detection2102.13086v2 · Xingyi Zhou, Vladlen Koltun, Philipp Krahenbuhl2021 · 2 citationsarXiv
Global Tracking Transformers2203.13250v2 · Xingyi Zhou, Tianwei Yin, Vladlen Koltun et al.2022 · 0 citationsarXiv
Multimodal Virtual Point 3D Detection2111.06881v1 · Tianwei Yin, Xingyi Zhou, Philipp Krahenbuhl2021 · 127 citationsarXiv
Center-based 3D Object Detection and Tracking2006.11275v2 · Tianwei Yin, Xingyi Zhou, Philipp Krahenbuhl2020 · 2,086 citationsProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2021
Learning to drive from a world on rails2105.00636v3 · Dian Chen, Vladlen Koltun, Philipp Krahenbuhl2021 · 103 citationsarXiv
Towards Long-Form Video Understanding2106.11310v1 · Chao-Yuan Wu, Philipp Krahenbuhl2021 · 120 citationsarXiv
Memory Optimization for Deep Networks2010.14501v3 · Aashaka Shah, Chao-Yuan Wu, Jayashree Mohan et al.2020 · 11 citationsarXiv
Probabilistic two-stage detection2103.07461v1 · Xingyi Zhou, Vladlen Koltun, Philipp Krahenbuhl2021 · 160 citationsarXiv
Tracking Objects as Points2004.01177v2 · Xingyi Zhou, Vladlen Koltun, Philipp Krahenbuhl2020 · 18 citationsarXiv
A Multigrid Method for Efficiently Training Video Models1912.00998v2 · Chao-Yuan Wu, Ross Girshick, Kaiming He et al.2019 · 88 citationsarXiv
Lossless Image Compression through Super-Resolution2004.02872v1 · Sheng Cao, Chao-Yuan Wu, Philipp Krahenbuhl2020 · 28 citationsarXiv
Learning by Cheating1912.12294v1 · Dian Chen, Brady Zhou, Vladlen Koltun et al.2019 · 3 citationsarXiv
Joint Monocular 3D Vehicle Detection and Tracking1811.10742v3 · Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang et al.2018 · 266 citationsarXiv
A Century of Portraits: A Visual Historical Record of American High School Yearbooks1511.02575v2 · Shiry Ginosar, Kate Rakelly, Sarah Sachs et al.2015 · 77 citationsarXiv
Monocular Plan View Networks for Autonomous Driving1905.06937v1 · Dequan Wang, Coline Devin, Qi-Zhi Cai et al.2019 · 77 citationsarXiv
Bottom-up Object Detection by Grouping Extreme and Center Points1901.08043v3 · Xingyi Zhou, Jiacheng Zhuo, Philipp Krahenbuhl2019 · 1,115 citationsarXiv
Long-Term Feature Banks for Detailed Video Understanding1812.05038v2 · Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan et al.2018 · 480 citationsarXiv
Assessing Generalization in Deep Reinforcement Learning1810.12282v2 · Charles Packer, Katelyn Gao, Jernej Kos et al.2018 · 114 citationsarXiv
Generative Visual Manipulation on the Natural Image Manifold1609.03552v3 · Jun-Yan Zhu, Philipp Krahenbuhl, Eli Shechtman et al.2016 · 1,294 citationsarXiv
Video Compression through Image Interpolation1804.06919v1 · Chao-Yuan Wu, Nayan Singhal, Philipp Krahenbuhl2018 · 334 citationsarXiv
Career total: 98 works. 62 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.