AT

Alexander Toshev

cs.CVcs.LGcs.ROcs.AIcs.CLstat.MLcs.GRcs.HCcs.NEcs.SY

On Valency

published · living versions
W_et6jfc9y·v1 · currentpublished
Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
with Oriol Vinyals, Samy Bengio, Dumitru Erhan
1 version

Preprints & journals

54 papers in the corpus · 2013–2026
Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models2607.16900v3 · Seanie Lee, Sanjoy Chowdhury, Chao Jiang et al.2026 · 0 citationsarXiv
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action2510.17790v3 · Yuhao Yang, Zhen Yang, Zi-Yi Dou et al.2025 · 0 citationsarXiv
Expanding LLM Agent Boundaries with Strategy-Guided Exploration2603.02045v1 · Andrew Szot, Michael Kirchhof, Omar Attia et al.2026 · 0 citationsarXiv
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning2510.02180v2 · Silvia Sapora, Devon Hjelm, Alexander Toshev et al.2025 · 0 citationsarXiv
Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality2503.07879v2 · Alex Fang, Hadi Pouransari, Matt Jordan et al.2025 · 0 citationsarXiv
Scaling Synthetic Task Generation for Agents via Exploration2509.25047v1 · Ram Ramrakhya, Andrew Szot, Omar Attia et al.2025 · 0 citationsarXiv
MobileCLIP2: Improving Multi-Modal Reinforced Training2508.20691v1 · Fartash Faghri, Pavan Kumar Anasosalu Vasu, Cem Koc et al.2025 · 0 citationsarXiv
DataComp-LM: In search of the next generation of training sets for language models2406.11794v4 · Jeffrey Li, Alex Fang, Georgios Smyrnis et al.2024 · 17 citationsarXiv
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models2412.09648v1 · Kevin Miao, Harsh Agrawal, Qihang Zhang et al.2024 · 0 citationsarXiv
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons2412.08442v1 · Andrew Szot, Bogdan Mazoure, Omar Attia et al.2024 · 8 citationsarXiv
Grounding Multimodal Large Language Models in Actions2406.07904v2 · Andrew Szot, Bogdan Mazoure, Harsh Agrawal et al.2024 · 1 citationarXiv
World-consistent Video Diffusion with Explicit 3D Modeling2412.01821v1 · Qihang Zhang, Shuangfei Zhai, Miguel Angel Bautista et al.2024 · 9 citationsarXiv
Multimodal Autoregressive Pre-training of Large Vision Encoders2411.14402v1 · Enrico Fini, Mustafa Shukor, Xiujun Li et al.2024 · 21 citationsarXiv
On the Modeling Capabilities of Large Language Models for Sequential Decision Making2410.05656v1 · Martin Klissarov, Devon Hjelm, Alexander Toshev et al.2024 · 0 citationsarXiv
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation2311.16201v2 · Yuhui Zhang, Brandon McKinzie, Zhe Gan et al.2023 · 1 citationarXiv
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training2403.09611v4 · Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier et al.2024 · 11 citationsarXiv
Large Language Models as Generalizable Policies for Embodied Tasks2310.17722v2 · Andrew Szot, Max Schwarzer, Harsh Agrawal et al.2023 · 8 citationsarXiv
Scalable Pre-training of Large Autoregressive Image Models2401.08541v1 · Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai et al.2024 · 6 citationsarXiv
Data Filtering Networks2309.17425v3 · Alex Fang, Albin Madappally Jose, Amit Jain et al.2023 · 15 citationsarXiv
Principles and Guidelines for Evaluating Social Robot Navigation Algorithms2306.16740v4 · Anthony Francis, Claudia P'erez-D'Arpino, Chengshu Li et al.2023 · 91 citationsarXiv
Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts2309.04354v1 · Erik Daxberger, Floris Weers, Bowen Zhang et al.2023 · 3 citationsarXiv
Perceptual Grouping in Contrastive Vision-Language Models2210.09996v3 · Kanchana Ranasinghe, Brandon McKinzie, Sachin Ravi et al.2022 · 40 citationsarXiv
Value function estimation using conditional diffusion models for control2306.07290v1 · Bogdan Mazoure, Walter Talbott, Miguel Angel Bautista et al.2023 · 0 citationsarXiv
On Robustness in Multimodal Learning2304.04385v2 · Brandon McKinzie, Joseph Cheng, Vaishaal Shankar et al.2023 · 0 citationsarXiv
STAIR: Learning Sparse Text and Image Representation in Grounded Tokens2301.13081v2 · Chen Chen, Bowen Zhang, Liangliang Cao et al.2023 · 17 citationsarXiv
Retrospectives on the Embodied AI Workshop2210.06849v3 · Matt Deitke, Dhruv Batra, Yonatan Bisk et al.2022 · 19 citationsarXiv
Gesture2Path: Imitation Learning for Gesture-aware Navigation2209.09375v1 · Catie Cuan, Edward Lee, Emre Fisher et al.2022 · 2 citationsarXiv
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2204.01691v2 · Michael Ahn, Anthony Brohan, Noah Brown et al.2022 · 518 citationsarXiv
GAUDI: A Neural Architect for Immersive 3D Scene Generation2207.13751v1 · Miguel Angel Bautista, Pengsheng Guo, Samira Abnar et al.2022 · 58 citationsarXiv
Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation2203.15041v2 · Haresh Karnan, Anirudh Nair, Xuesu Xiao et al.2022 · 119 citationsRobotics and Automation Letters (RA-L) 2022
A Protocol for Validating Social Navigation Policies2204.05443v2 · Soren Pirk, Edward Lee, Xuesu Xiao et al.2022 · 6 citationsarXiv
Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning2111.03189v2 · Dhruv Shah, Peng Xu, Yao Lu et al.2021 · 11 citationsarXiv
ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation2008.07792v2 · Fei Xia, Chengshu Li, Roberto Mart'in-Mart'in et al.2020 · 36 citationsarXiv
Modeling Long-horizon Tasks as Sequential Interaction Landscapes2006.04843v2 · Soren Pirk, Karol Hausman, Alexander Toshev et al.2020 · 13 citationsarXiv
ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects2006.13171v2 · Dhruv Batra, Aaron Gokaslan, Aniruddha Kembhavi et al.2020 · 129 citationsarXiv
Adversarial Generative Grammars for Human Activity Prediction2008.04888v2 · AJ Piergiovanni, Anelia Angelova, Alexander Toshev et al.2020 · 25 citationsarXiv
Learning Object-conditioned Exploration using Distributed Soft Actor Critic2007.14545v2 · Ayzaan Wahid, Austin Stone, Kevin Chen et al.2020 · 13 citationsarXiv
Long Range Neural Navigation Policies for the Real World1903.09870v2 · Ayzaan Wahid, Alexander Toshev, Marek Fiser et al.2019 · 17 citationsarXiv
Evolving Space-Time Neural Architectures for Videos1811.10636v2 · AJ Piergiovanni, Anelia Angelova, Alexander Toshev et al.2018 · 81 citationsICCV 2019
Visual Representations for Semantic Target Driven Navigation1805.06066v3 · Arsalan Mousavian, Alexander Toshev, Marek Fiser et al.2018 · 242 citationsarXiv
Scene Memory Transformer for Embodied Agents in Long-Horizon Tasks1903.03878v1 · Kuan Fang, Alexander Toshev, Li Fei-Fei et al.2019 · 198 citationsarXiv
Self-supervisory Signals for Object Discovery and Detection1806.03370v1 · Etienne Pot, Alexander Toshev, Jana Kosecka2018 · 8 citationsarXiv
Sim2Real View Invariant Visual Servoing by Recurrent Control1712.07642v1 · Fereshteh Sadeghi, Alexander Toshev, Eric Jang et al.2017 · 43 citationsarXiv
No Fuss Distance Metric Learning using Proxies1703.07464v3 · Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung et al.2017 · 698 citationsarXiv
Towards Accurate Multi-person Pose Estimation in the Wild1701.01779v2 · George Papandreou, Tyler Zhu, Nori Kanazawa et al.2017 · 977 citationsarXiv
DeepPose: Human Pose Estimation via Deep Neural Networks1312.4659v3 · Alexander Toshev, Christian Szegedy2013 · 3,357 citationsarXivon Valency
Chained Predictions Using Convolutional Neural Networks1605.02346v2 · Georgia Gkioxari, Alexander Toshev, Navdeep Jaitly2016 · 211 citationsarXiv
The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition1511.06789v3 · Jonathan Krause, Benjamin Sapp, Andrew Howard et al.2015 · 327 citationsarXiv
Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge1609.06647v1 · Oriol Vinyals, Alexander Toshev, Samy Bengio et al.2016 · 931 citationsIEEE Transactions on Pattern Analysis and Machine Intelligence ( Volume: PP, Issue: 99 , July 2016 )on Valency
Generation and Comprehension of Unambiguous Object Descriptions1511.02283v3 · Junhua Mao, Jonathan Huang, Alexander Toshev et al.2015 · 1,280 citationsarXiv
Career total: 95 works. 54 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.