EW

Eric Wallace

cs.CLcs.LGcs.AIcs.CRcs.CYstat.MLcs.CVcs.SDeess.AS

On Valency

published · living versions
W_jsb5733n·v1 · currentpublished
Deliberative Alignment: Reasoning Enables Safer Language Models
with Melody Y. Guan, Manas Joglekar, Saachi Jain, Boaz Barak +10
1 version

Preprints & journals

26 papers in the corpus · 2018–2024
Deliberative Alignment: Reasoning Enables Safer Language Models2412.16339v2 · Melody Y. Guan, Manas Joglekar, Eric Wallace et al.2024 · 34 citationsarXivon Valency
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
What Evidence Do Language Models Find Convincing?2402.11782v2 · Alexander Wan, Eric Wallace, Dan Klein2024 · 9 citationsarXiv
Stealing Part of a Production Language Model2403.06634v2 · Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham et al.2024 · 7 citationsarXiv
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions2404.13208v1 · Eric Wallace, Kai Xiao, Reimar Leike et al.2024 · 10 citationsarXiv
Analyzing Dynamic Adversarial Training Data in the Limit2110.08514v2 · Eric Wallace, Adina Williams, Robin Jia et al.2021 · 12 citationsarXiv
Pathologies of Neural Models Make Interpretations Difficult1804.07781v3 · Shi Feng, Eric Wallace, Alvin Grissom II et al.2018 · 266 citationsarXiv
Cutting Down on Prompts and Parameters: Simple Few-Shot Learning with Language Models2106.13353v2 · Robert L. Logan IV, Ivana Balazevi'c, Eric Wallace et al.2021 · 135 citationsarXiv
Extracting Training Data from Large Language Models2012.07805v2 · Nicholas Carlini, Florian Tramer, Eric Wallace et al.2020 · 273 citationsarXiv
Calibrate Before Use: Improving Few-Shot Performance of Language Models2102.09690v2 · Tony Z. Zhao, Eric Wallace, Shi Feng et al.2021 · 72 citationsarXiv
Detoxifying Language Models Risks Marginalizing Minority Voices2104.06390v1 · Albert Xu, Eshaan Pathak, Eric Wallace et al.2021 · 69 citationsarXiv
Concealed Data Poisoning Attacks on NLP Models2010.12563v2 · Eric Wallace, Tony Z. Zhao, Shi Feng et al.2020 · 108 citationsarXiv
Universal Adversarial Triggers for Attacking and Analyzing NLP1908.07125v3 · Eric Wallace, Shi Feng, Nikhil Kandpal et al.2019 · 652 citationsarXiv
Imitation Attacks and Defenses for Black-box Machine Translation Systems2004.15015v3 · Eric Wallace, Mitchell Stern, Dawn Song2020 · 80 citationsarXiv
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts2010.15980v2 · Taylor Shin, Yasaman Razeghi, Robert L. Logan IV et al.2020 · 1,269 citationsarXiv
Gradient-based Analysis of NLP Models is Manipulable2010.05419v1 · Junlin Wang, Jens Tuyls, Eric Wallace et al.2020 · 46 citationsarXiv
Evaluating Models' Local Decision Boundaries via Contrast Sets2004.02709v2 · Matt Gardner, Yoav Artzi, Victoria Basmova et al.2020 · 285 citationsarXiv
Pretrained Transformers Improve Out-of-Distribution Robustness2004.06100v2 · Dan Hendrycks, Xiaoyuan Liu, Eric Wallace et al.2020 · 316 citationsarXiv
AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models1909.09251v1 · Eric Wallace, Jens Tuyls, Junlin Wang et al.2019 · 138 citationsarXiv
Do NLP Models Know Numbers? Probing Numeracy in Embeddings1909.07940v2 · Eric Wallace, Yizhong Wang, Sujian Li et al.2019 · 230 citationsarXiv
Trick Me If You Can: Human-in-the-loop Generation of Adversarial Examples for Question Answering1809.02701v4 · Eric Wallace, Pedro Rodriguez, Shi Feng et al.2018 · 138 citationsarXiv
Misleading Failures of Partial-input Baselines1905.05778v3 · Shi Feng, Eric Wallace, Jordan Boyd-Graber2019 · 35 citationsarXiv
Compositional Questions Do Not Necessitate Multi-hop Reasoning1906.02900v1 · Sewon Min, Eric Wallace, Sameer Singh et al.2019 · 151 citationsarXiv
Interpreting Neural Networks With Nearest Neighbors1809.02847v2 · Eric Wallace, Shi Feng, Jordan Boyd-Graber2018 · 66 citationsarXiv
Career total: 48 works. 26 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.