NC

Nicholas Carlini

cs.LGcs.CRcs.AIcs.CVstat.MLcs.CLcs.SEcs.CYcs.DBcs.DC

On Valency

published · living versions
W_u7fa5rax·v1 · currentpublished
MLSys: The New Frontier of Machine Learning Systems
with Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G. Andersen +64
1 version

Preprints & journals

106 papers in the corpus · 2016–2026
Large-scale online deanonymization with LLMs2602.16800v2 · Simon Lermen, Daniel Paleka, Joshua Swanson et al.2026 · 1 citation35th USENIX Security Symposium (USENIX Security 26), pp. 1947-1966 (2026)
CryptanalysisBench: Can LLMs do Cryptanalysis?2607.18538v2 · Lukas Fluri, Avital Shafran, Nicholas Carlini et al.2026 · 0 citationsarXiv
Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains2604.02343v2 · Roy Rinberg, Annabelle Michael Carrell, Simon Henniger et al.2026 · 0 citationsarXiv
Position: Adversarial ML for LLMs Is Not Making Any Progress2502.02260v2 · Javier Rando, Jie Zhang, Nicholas Carlini et al.2025 · 1 citationarXiv
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?2605.11086v1 · Zhun Wang, Nico Schiller, Hongwei Li et al.2026 · 3 citationsarXiv
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces2601.11868v1 · Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini et al.2026 · 0 citationsarXiv
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases2510.20270v1 · Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini2025 · 0 citationsarXiv
Detecting Adversarial Fine-tuning with Auditing Agents2510.16255v1 · Sarah Egler, John Schulman, Nicholas Carlini2025 · 0 citationsarXiv
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples2510.07192v1 · Alexandra Souly, Javier Rando, Ed Chapman et al.2025 · 7 citationsarXiv
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks2510.01676v1 · Milad Nasr, Yanick Fratantonio, Luca Invernizzi et al.2025 · 0 citationsarXiv
Defending Against Prompt Injection With a Few DefensiveTokens2507.07974v2 · Sizhe Chen, Yizhu Wang, Nicholas Carlini et al.2025 · 17 citationsarXiv
SoK: Watermarking for AI-Generated Content2411.18479v3 · Xuandong Zhao, Sam Gunn, Miranda Christ et al.2024 · 21 citationsarXiv
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising2403.11981v2 · Sanghyun Hong, Nicholas Carlini, Alexey Kurakin2024 · 0 citationsarXiv
LLMs unlock new paths to monetizing exploits2505.11449v1 · Nicholas Carlini, Milad Nasr, Edoardo Debenedetti et al.2025 · 0 citationsarXiv
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses2503.01811v1 · Nicholas Carlini, Javier Rando, Edoardo Debenedetti et al.2025 · 1 citationarXiv
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI2406.12027v2 · Robert Honig, Javier Rando, Nicholas Carlini et al.2024 · 3 citationsarXiv
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense2411.14834v2 · Jie Zhang, Christian Schlarmann, Kristina Nikoli'c et al.2024 · 0 citationsarXiv
On Evaluating the Durability of Safeguards for Open-Weight LLMs2412.07097v1 · Xiangyu Qi, Boyi Wei, Nicholas Carlini et al.2024 · 0 citationsarXiv
Query-Based Adversarial Prompt Generation2402.12329v2 · Jonathan Hayase, Ema Borevkovic, Nicholas Carlini et al.2024 · 6 citationsarXiv
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models2411.10242v1 · Michael Aerni, Javier Rando, Edoardo Debenedetti et al.2024 · 0 citationsarXiv
Stealing User Prompts from Mixture of Experts2410.22884v1 · Itay Yona, Ilia Shumailov, Jamie Hayes et al.2024 · 3 citationsarXiv
Remote Timing Attacks on Efficient Language Model Inference2410.17175v1 · Nicholas Carlini, Milad Nasr2024 · 1 citationarXiv
Persistent Pre-Training Poisoning of LLMs2410.13722v1 · Yiming Zhang, Javier Rando, Ivan Evtimov et al.2024 · 1 citationarXiv
Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label Setting2410.05750v1 · Nicholas Carlini, Jorge Ch'avez-Saab, Anna Hambitzer et al.2024 · 10 citationsarXiv
Forcing Diffuse Distributions out of Language Models2404.10859v2 · Yiming Zhang, Avi Schwarzschild, Nicholas Carlini et al.2024 · 0 citationsarXiv
Privacy Side Channels in Machine Learning Systems2309.05610v2 · Edoardo Debenedetti, Giorgio Severi, Nicholas Carlini et al.2023 · 5 citationsarXiv
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining2212.06470v3 · Florian Tramer, Gautam Kamath, Nicholas Carlini2022 · 9 citationsarXiv
Stealing Part of a Production Language Model2403.06634v2 · Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham et al.2024 · 7 citationsarXiv
Poisoning Web-Scale Training Datasets is Practical2302.10149v2 · Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo et al.2023 · 132 citationsarXiv
Are aligned neural networks adversarially aligned?2306.15447v2 · Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo et al.2023 · 50 citationsarXiv
Initialization Matters for Adversarial Transfer Learning2312.05716v2 · Andong Hua, Jindong Gu, Zhiyu Xue et al.2023 · 6 citationsarXiv
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models2404.01231v1 · Yuxin Wen, Leo Marchyok, Sanghyun Hong et al.2024 · 4 citationsarXiv
Evading Black-box Classifiers Without Breaking Eggs2306.02895v2 · Edoardo Debenedetti, Nicholas Carlini, Florian Tramer2023 · 4 citationsarXiv
Identifying and Mitigating the Security Risks of Generative AI2308.14840v4 · Clark Barrett, Brad Boyd, Elie Burzstein et al.2023 · 63 citationsFoundations and Trends in Privacy and Security 6 (2023) 1-52
Universal and Transferable Adversarial Attacks on Aligned Language Models2307.15043v2 · Andy Zou, Zifan Wang, Nicholas Carlini et al.2023 · 196 citationsarXiv
Report of the 1st Workshop on Generative AI and Law2311.06477v3 · A. Feder Cooper, Katherine Lee, James Grimmelmann et al.2023 · 6 citationsarXiv
Scalable Extraction of Training Data from (Production) Language Models2311.17035v1 · Milad Nasr, Nicholas Carlini, Jonathan Hayase et al.2023 · 83 citationsarXiv
Effective Robustness against Natural Distribution Shifts for Models with Different Training Data2302.01381v2 · Zhouxing Shi, Nicholas Carlini, Ananth Balashankar et al.2023 · 0 citationsarXiv
Counterfactual Memorization in Neural Language Models2112.12938v2 · Chiyuan Zhang, Daphne Ippolito, Katherine Lee et al.2021 · 50 citationsarXiv
Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy2210.17546v3 · Daphne Ippolito, Florian Tramer, Milad Nasr et al.2022 · 19 citationsarXiv
Reverse-Engineering Decoding Strategies Given Blackbox Access to a Language Generation System2309.04858v1 · Daphne Ippolito, Nicholas Carlini, Katherine Lee et al.2023 · 5 citationsarXiv
Backdoor Attacks for In-Context Learning with Language Models2307.14692v1 · Nikhil Kandpal, Matthew Jagielski, Florian Tramer et al.2023 · 9 citationsarXiv
A LLM Assisted Exploitation of AI-Guardian2307.15008v1 · Nicholas Carlini2023 · 2 citationsarXiv
Preprocessors Matter! Realistic Decision-Based Attacks on Machine Learning Systems2210.03297v2 · Chawin Sitawarin, Florian Tramer, Nicholas Carlini2022 · 4 citationsarXiv
Measuring Forgetting of Memorized Training Examples2207.00099v2 · Matthew Jagielski, Om Thakkar, Florian Tramer et al.2022 · 15 citationsarXiv
Part-Based Models Improve Adversarial Robustness2209.09117v2 · Chawin Sitawarin, Kornrapat Pongmala, Yizheng Chen et al.2022 · 2 citationsarXiv
Students Parrot Their Teachers: Membership Inference on Model Distillation2303.03446v1 · Matthew Jagielski, Milad Nasr, Christopher Choquette-Choo et al.2023 · 8 citationsarXiv
Career total: 160 works. 106 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.