AB
Alex Beutel
cs.LGcs.AIstat.MLcs.CLcs.CYcs.IRcs.CVcs.SIcs.CRcs.DB
On Valency
published · living versionsW_jsb5733n·v1 · currentpublished
Deliberative Alignment: Reasoning Enables Safer Language Models
with Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain +10
1 version
Preprints & journals
54 papers in the corpus · 2014–2026Artificial Intelligence-Based Simulation to Improve Code Status Discussions Among Internal Medicine Residents: A Pilot Randomized Trial.42690594 · Lichter, Jessica R, Beutel, Alex, Trottier, Dana et al.2026 · 0 citationsJournal of general internal medicine. 2026
OpenAI GPT-5 System Card2601.03267v2 · Aaditya Singh, Adam Fry, Adam Perelman et al.2025 · 18 citationsarXiv
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs2603.10521v1 · Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu et al.2026 · 0 citationsarXiv
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training2508.09224v1 · Yuan Yuan, Tina Sriskandarajah, Anna-Luisa Brakman et al.2025 · 5 citationsarXiv
First-Person Fairness in Chatbots2410.19803v2 · Tyna Eloundou, Alex Beutel, David G. Robinson et al.2024 · 3 citationsarXiv
Deliberative Alignment: Reasoning Enables Safer Language Models2412.16339v2 · Melody Y. Guan, Manas Joglekar, Eric Wallace et al.2024 · 34 citationsarXivon Valency
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning2412.18693v1 · Alex Beutel, Kai Xiao, Johannes Heidecke et al.2024 · 1 citationarXiv
Rule Based Rewards for Language Model Safety2411.01111v1 · Tong Mu, Alec Helyar, Johannes Heidecke et al.2024 · 10 citationsarXiv
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Controlled Decoding from Language Models2310.17022v3 · Sidharth Mudgal, Jong Lee, Harish Ganapathy et al.2023 · 2 citationsarXiv
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing2312.03867v2 · Lucas Monteiro Paes, Ananda Theertha Suresh, Alex Beutel et al.2023 · 3 citationsarXiv
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions2404.13208v1 · Eric Wallace, Kai Xiao, Reimar Leike et al.2024 · 10 citationsarXiv
Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks2310.16955v2 · Aradhana Sinha, Ananth Balashankar, Ahmad Beirami et al.2023 · 0 citationsTransactions on Machine Learning Research (2024)
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images2401.14322v1 · Hansa Srinivasan, Candice Schumann, Aradhana Sinha et al.2024 · 1 citationarXiv
Effective Robustness against Natural Distribution Shifts for Models with Different Training Data2302.01381v2 · Zhouxing Shi, Nicholas Carlini, Ananth Balashankar et al.2023 · 0 citationsarXiv
Improving Few-shot Generalization of Safety Classifiers via Data Augmented Parameter-Efficient Fine-Tuning2310.16959v1 · Ananth Balashankar, Xiao Ma, Aradhana Sinha et al.2023 · 0 citationsarXiv
Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting2310.16523v1 · Preethi Lahoti, Nicholas Blumm, Xiao Ma et al.2023 · 16 citationsarXiv
Learning from Negative User Feedback and Measuring Responsiveness for Sequential Recommenders2308.12256v1 · Yueqi Wang, Yoni Halpern, Shuo Chang et al.2023 · 12 citationsarXiv
Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification2307.05728v1 · James Atwood, Tina Tian, Ben Packer et al.2023 · 0 citationsarXiv
Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning2306.14308v1 · Xiao Ma, Swaroop Mishra, Ahmad Beirami et al.2023 · 3 citationsarXiv
Improving Classifier Robustness through Active Generation of Pairwise Counterfactuals2305.13535v1 · Ananth Balashankar, Xuezhi Wang, Yao Qin et al.2023 · 0 citationsarXiv
Towards Robust Prompts on Vision-Language Models2304.08479v1 · Jindong Gu, Ahmad Beirami, Xuezhi Wang et al.2023 · 2 citationsarXiv
Understanding and Improving Robustness of Vision Transformers through Patch-based Negative Augmentation2110.07858v2 · Yao Qin, Chiyuan Zhang, Ting Chen et al.2021 · 16 citationsarXiv
What Are Effective Labels for Augmented Data? Improving Calibration and Robustness with AutoLabel2302.11188v1 · Yao Qin, Xuezhi Wang, Balaji Lakshminarayanan et al.2023 · 4 citationsarXiv
Striving for data-model efficiency: Identifying data externalities on group performance2211.06348v1 · Esther Rolf, Ben Packer, Alex Beutel et al.2022 · 0 citationsarXiv
A Human-ML Collaboration Framework for Improving Video Content Reviews2210.09500v1 · Meghana Deodhar, Xiao Ma, Yixin Cai et al.2022 · 0 citationsarXiv
Simpson's Paradox in Recommender Fairness: Reconciling differences between per-user and aggregated evaluations2210.07755v1 · Flavien Prost, Ben Packer, Jilin Chen et al.2022 · 0 citationsarXiv
Flexible text generation for counterfactual fairness probing2206.13757v1 · Zee Fryer, Vera Axelrod, Ben Packer et al.2022 · 9 citationsarXiv
Top-K Off-Policy Correction for a REINFORCE Recommender System1812.02353v3 · Minmin Chen, Alex Beutel, Paul Covington et al.2018 · 414 citationsarXiv
Improving Calibration through the Relationship with Adversarial Robustness2006.16375v2 · Yao Qin, Xuezhi Wang, Alex Beutel et al.2020 · 8 citationsarXiv
Measuring Model Fairness under Noisy Covariates: A Theoretical Perspective2105.09985v1 · Flavien Prost, Pranjal Awasthi, Nick Blumm et al.2021 · 6 citationsarXiv
Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities2105.02377v1 · Ruohan Zhan, Konstantina Christakopoulou, Ya Le et al.2021 · 2 citationsarXiv
Measuring and Reducing Gendered Correlations in Pre-trained Models2010.06032v2 · Kellie Webster, Xuezhi Wang, Ian Tenney et al.2020 · 106 citationsarXiv
Evaluating Fairness of Machine Learning Models Under Uncertain and Incomplete Information2102.08410v1 · Pranjal Awasthi, Alex Beutel, Matthaeus Kleindessner et al.2021 · 42 citationsarXiv
Practical Compositional Fairness: Understanding Fairness in Multi-Component Recommender Systems1911.01916v4 · Xuezhi Wang, Nithum Thain, Anu Sinha et al.2019 · 17 citationsarXiv
Measuring Recommender System Effects with Simulated Users2101.04526v1 · Sirui Yao, Yoni Halpern, Nithum Thain et al.2021 · 20 citationsarXiv
Learned Indexes for a Google-scale Disk-based Database2012.12501v1 · Hussam Abu-Libdeh, Deniz Alt\inbuken, Alex Beutel et al.2020 · 14 citationsarXiv
Underspecification Presents Challenges for Credibility in Modern Machine Learning2011.03395v2 · Alexander D'Amour, Katherine Heller, Dan Moldovan et al.2020 · 420 citationsarXiv
Fairness without Demographics through Adversarially Reweighted Learning2006.13114v3 · Preethi Lahoti, Alex Beutel, Jilin Chen et al.2020 · 46 citationsarXiv
CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text Generation2010.02338v1 · Tianlu Wang, Xuezhi Wang, Yao Qin et al.2020 · 70 citationsarXiv
Transfer of Machine Learning Fairness across Domains1906.09688v3 · Candice Schumann, Xuezhi Wang, Alex Beutel et al.2019 · 27 citationsarXiv
Toward a better trade-off between performance and fairness with kernel-based distribution matching1910.11779v1 · Flavien Prost, Hai Qian, Qiuwen Chen et al.2019 · 13 citationsarXiv
Fairness in Recommendation Ranking through Pairwise Comparisons1903.00780v1 · Alex Beutel, Jilin Chen, Tulsee Doshi et al.2019 · 366 citationsarXiv
Counterfactual Fairness in Text Classification through Robustness1809.10610v2 · Sahaj Garg, Vincent Perot, Nicole Limtiaco et al.2018 · 211 citationsarXiv
Putting Fairness Principles into Practice: Challenges, Metrics, and Improvements1901.04562v1 · Alex Beutel, Jilin Chen, Tulsee Doshi et al.2019 · 25 citationsarXiv
BIRDNEST: Bayesian Inference for Ratings-Fraud Detection1511.06030v2 · Bryan Hooi, Neil Shah, Alex Beutel et al.2015 · 108 citationsarXiv
The Case for Learned Index Structures1712.01208v3 · Tim Kraska, Alex Beutel, Ed H. Chi et al.2017 · 936 citationsarXiv
The Many Faces of Link Fraud1704.01420v3 · Neil Shah, Hemank Lamba, Alex Beutel et al.2017 · 34 citationsarXiv
Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations1707.00075v2 · Alex Beutel, Jilin Chen, Zhe Zhao et al.2017 · 292 citationsarXiv
Career total: 116 works. 54 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.