cs.LGcs.AIcs.CLcs.CRcs.CVcs.CYstat.ML

On Valency

published · living versions
W_aevsngy8·v1 · currentpublished
Trading Inference-Time Compute for Adversarial Robustness
with Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak, Stephanie Lin +6
1 version

Preprints & journals

17 papers in the corpus · 2017–2026
GPT-Red: Automated Red Teaming via Self-Play at Scale2607.26115v1 · Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal et al.2026 · 0 citationsarXiv
OpenAI GPT-5 System Card2601.03267v2 · Aaditya Singh, Adam Fry, Adam Perelman et al.2025 · 18 citationsarXiv
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs2603.10521v1 · Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu et al.2026 · 0 citationsarXiv
Trading Inference-Time Compute for Adversarial Robustness2501.18841v1 · Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak et al.2025 · 0 citationsarXivon Valency
Deliberative Alignment: Reasoning Enables Safer Language Models2412.16339v2 · Melody Y. Guan, Manas Joglekar, Eric Wallace et al.2024 · 34 citationsarXivon Valency
Human Action Anticipation: A Survey2410.14045v1 · Bolin Lai, Sam Toyer, Tushar Nagarajan et al.2024 · 0 citationsarXiv
A StrongREJECT for Empty Jailbreaks2402.10260v2 · Alexandra Souly, Qingyuan Lu, Dillon Bowen et al.2024 · 23 citationsarXiv
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game2311.01011v1 · Sam Toyer, Olivia Watkins, Ethan Adrian Mendes et al.2023 · 8 citationsarXiv
imitation: Clean Imitation Learning Implementations2211.11972v1 · Adam Gleave, Mohammad Taufeeque, Juan Rocamonde et al.2022 · 9 citationsarXiv
An Empirical Investigation of Representation Learning for Imitation2205.07886v1 · Xin Chen, Sam Toyer, Cody Wild et al.2022 · 7 citationsarXiv
DERAIL: Diagnostic Environments for Reward And Imitation Learning2012.01365v1 · Pedro Freire, Adam Gleave, Sam Toyer et al.2020 · 1 citationarXiv
The MAGICAL Benchmark for Robust Imitation2011.00401v1 · Sam Toyer, Rohin Shah, Andrew Critch et al.2020 · 7 citationsarXiv
Action Schema Networks: Generalised Policies with Deep Learning1709.04271v2 · Sam Toyer, Felipe Trevizan, Sylvie Thi'ebaux et al.2017 · 73 citationsarXiv
Human Pose Forecasting via Deep Markov Models1707.09240v2 · Sam Toyer, Anoop Cherian, Tengda Han et al.2017 · 45 citationsarXiv
Career total: 24 works. 17 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.