DS
Dawn Song
cs.CRcs.LGcs.AIcs.CLstat.MLcs.CVcs.CYcs.SEcs.DCcs.PL
On Valency
published · living versionsW_u7fa5rax·v1 · currentpublished
MLSys: The New Frontier of Machine Learning Systems
with Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G. Andersen +64
1 version
Preprints & journals
190 papers in the corpus · 2009–2026Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs2502.10673v2 · Yepeng Liu, Xuandong Zhao, Dawn Song et al.2025 · 0 citationsarXiv
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents2602.08235v2 · Jaylen Jones, Zhehao Zhang, Yuting Ning et al.2026 · 0 citationsarXiv
CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence2510.11974v2 · Yutong Cheng, Yang Liu, Changze Li et al.2025 · 1 citationarXiv
Learning to Reason without External Rewards2505.19590v5 · Xuandong Zhao, Zhewei Kang, Aosong Feng et al.2025 · 0 citationsarXiv
Progent: Securing AI Agents with Privilege Control2504.11703v3 · Tianneng Shi, Jingxuan He, Zhun Wang et al.2025 · 0 citationsarXiv
RepIt: Steering Language Models with Concept-Specific Refusal Vectors2509.13281v5 · Vincent Siu, Nathan W. Henry, Nicholas Crispino et al.2025 · 0 citationsarXiv
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption2510.18333v2 · Yepeng Liu, Xuandong Zhao, Dawn Song et al.2025 · 0 citationsarXiv
In-Context Watermarks for Large Language Models2505.16934v2 · Yepeng Liu, Xuandong Zhao, Christopher Kruegel et al.2025 · 0 citationsarXiv
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence2505.11611v3 · Bofan Gong, Shiyang Lai, James Evans et al.2025 · 0 citationsICLR 2026
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents2506.14205v2 · Jingxu Xie, Dylan Xu, Xuandong Zhao et al.2025 · 0 citationsarXiv
TxRay: Agentic Postmortem of Live Blockchain Attacks2602.01317v5 · Ziyue Wang, Jiangshan Yu, Kaihua Qin et al.2026 · 0 citationsarXiv
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle2601.20882v1 · Yuheng Tang, Kaijie Zhu, Bonan Ruan et al.2026 · 0 citationsarXiv
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?2504.11741v2 · Yiyou Sun, Georgia Zhou, Haoyue Bai et al.2025 · 0 citationsarXiv
Scalable Best-of-N Selection for Large Language Models via Self-Certainty2502.18581v3 · Zhewei Kang, Xuandong Zhao, Dawn Song2025 · 1 citationarXiv
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation2507.05578v2 · Alexander Xiong, Xuandong Zhao, Aneesh Pappu et al.2025 · 1 citationarXiv
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems2505.15216v3 · Andy K. Zhang, Joey Ji, Celeste Menders et al.2025 · 1 citationarXiv
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning2505.16186v2 · Kaiwen Zhou, Xuandong Zhao, Gaowen Liu et al.2025 · 0 citationsarXiv
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents2510.02609v2 · Chengquan Guo, Chulin Xie, Yu Yang et al.2025 · 0 citationsarXiv
Scaling Agent Learning via Experience Synthesis2511.03773v2 · Zhaorun Chen, Zhuokai Zhao, Kai Zhang et al.2025 · 0 citationsarXiv
An Illusion of Progress? Assessing the Current State of Web Agents2504.01382v4 · Tianci Xue, Weijian Qi, Tianneng Shi et al.2025 · 0 citationsCOLM 2025
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?2509.21016v2 · Yiyou Sun, Yuhan Cao, Pohao Huang et al.2025 · 0 citationsarXiv
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI2410.11096v2 · Yuzhou Nie, Zhun Wang, Yu Yang et al.2024 · 1 citationarXiv
Wrangling Entropy: Next-Generation Multi-Factor Key Derivation, Credential Hashing, and Credential Generation Functions2509.05893v1 · Colin Roberts, Vivek Nair, Dawn Song2025 · 0 citationsarXiv
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage2412.05734v2 · Yuzhou Nie, Zhun Wang, Ye Yu et al.2024 · 1 citationarXiv
Advancing Science- and Evidence-based AI Policy2508.02748v1 · Rishi Bommasani, Sanjeev Arora, Jennifer Chayes et al.2025 · 22 citationsarXiv
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis2410.01817v2 · Tanusree Sharma, Yujin Potter, Zachary Kilhoffer et al.2024 · 0 citationsarXiv
WebGuard: Building a Generalizable Guardrail for Web Agents2507.14293v1 · Boyuan Zheng, Zeyi Liao, Scott Salisbury et al.2025 · 0 citationsarXiv
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models2507.07484v1 · Kaiqu Liang, Haimin Hu, Xuandong Zhao et al.2025 · 3 citationsarXiv
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models2411.10557v3 · Jianhong Tu, Zhuohao Ni, Nicholas Crispino et al.2024 · 1 citationarXiv
Decompiling Smart Contracts with a Large Language Model2506.19624v1 · Isaac David, Liyi Zhou, Dawn Song et al.2025 · 0 citationsarXiv
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents2505.05849v4 · Zhun Wang, Vincent Siu, Zhe Ye et al.2025 · 0 citationsarXiv
SoK: Watermarking for AI-Generated Content2411.18479v3 · Xuandong Zhao, Sam Gunn, Miranda Christ et al.2024 · 21 citationsarXiv
COSMIC: Generalized Refusal Direction Identification in LLM Activations2506.00085v1 · Vincent Siu, Nicholas Crispino, Zihao Yu et al.2025 · 0 citationsarXiv
A Critical Evaluation of Defenses against Prompt Injection Attacks2505.18333v1 · Yuqi Jia, Zedian Shao, Yupei Liu et al.2025 · 1 citationarXiv
Type-Constrained Code Generation with Language Models2504.09246v2 · Niels Mundler, Jingxuan He, Hao Wang et al.2025 · 10 citationsarXiv
SoK: The Gap Between Data Rights Ideals and Reality2312.01511v2 · Yujin Potter, Ella Corren, Gonzalo Munilla Garrido et al.2023 · 1 citationarXiv
An Undetectable Watermark for Generative Image Models2410.07369v4 · Sam Gunn, Xuandong Zhao, Dawn Song2024 · 0 citationsarXiv
CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models2410.21060v2 · Yutong Cheng, Osama Bajaber, Saimon Amanuel Tsegai et al.2024 · 15 citationsarXiv
Enhancing Smart Contract Security Analysis with Execution Property Graphs2305.14046v3 · Kaihua Qin, Zhe Ye, Zhun Wang et al.2023 · 4 citationsarXiv
International Scientific Report on the Safety of Advanced AI (Interim Report)2412.05282v2 · Yoshua Bengio, Soren Mindermann, Daniel Privitera et al.2024 · 7 citationsarXiv
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models2503.14827v1 · Chejian Xu, Jiawei Zhang, Zhaorun Chen et al.2025 · 0 citationsarXiv
Representation Engineering: A Top-Down Approach to AI Transparency2310.01405v4 · Andy Zou, Long Phan, Sarah Chen et al.2023 · 65 citationsarXiv
DeServe: Towards Affordable Offline LLM Inference via Decentralization2501.14784v1 · Linyu Wu, Xiaoyuan Liu, Tianneng Shi et al.2025 · 0 citationsarXiv
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification2405.00253v4 · Yuchen Tian, Weixiang Yan, Qian Yang et al.2024 · 22 citationsarXiv
Formal Mathematical Reasoning: A New Frontier in AI2412.16075v1 · Kaiyu Yang, Gabriel Poesia, Jingxuan He et al.2024 · 3 citationsarXiv
RedCode: Risky Code Execution and Generation Benchmark for Code Agents2411.07781v1 · Chengquan Guo, Xun Liu, Chulin Xie et al.2024 · 11 citationsarXiv
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters2410.24190v3 · Yujin Potter, Shiyang Lai, Junsol Kim et al.2024 · 18 citationsarXiv
ThreatKG: An AI-Powered System for Automated Open-Source Cyber Threat Intelligence Gathering and Management2212.10388v2 · Peng Gao, Xiaoyuan Liu, Edward Choi et al.2022 · 27 citationsarXiv
LLM-PBE: Assessing Data Privacy in Large Language Models2408.12787v2 · Qinbin Li, Junyuan Hong, Chulin Xie et al.2024 · 39 citationsarXiv
Agent Instructs Large Language Models to be General Zero-Shot Reasoners2310.03710v2 · Nicholas Crispino, Kyle Montgomery, Fankun Zeng et al.2023 · 3 citationsarXiv
Career total: 528 works. 190 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.