AD
Allan Dafoe
cs.AIcs.CYcs.LGcs.CRcs.MAecon.GNq-fin.ECstat.MEAccess to InformationAgaonidae
On Valency
published · living versionsW_y3274kcu·v1 · currentpublished
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
with Miles Brundage, Shahar Avin, Jack Clark, Helen Toner +21
1 version
Preprints & journals
35 papers in the corpus · 2015–2026From AGI to ASI2606.12683v2 · Tim Genewein, Matija Franklin, Alexander Lerchner et al.2026 · 0 citationsarXiv
Comprehensive AI governance requires addressing non-model gains2606.00047v1 · Arthur Goemans, Dan Altman, Noemi Dreksler et al.2026 · 0 citationsarXiv
Measuring Progress Toward AGI: A Cognitive Framework2605.28405v1 · Ryan Burnell, Yumeya Yamamori, Orhan Firat et al.2026 · 0 citationsarXiv
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety2507.11473v2 · Tomek Korbak, Mikita Balesni, Elizabeth Barnes et al.2025 · 2 citationsarXiv
Levels of AGI for Operationalizing Progress on the Path to AGI2311.02462v5 · Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel et al.2023 · 53 citationsProceedings of ICML 2024
Evaluating Frontier Models for Stealth and Situational Awareness2505.01420v4 · Mary Phuong, Roland S. Zimmermann, Ziyue Wang et al.2025 · 0 citationsarXiv
Gemini: A Family of Highly Capable Multimodal Models2312.11805v5 · Gemini Team Google: Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et al.2023 · 838 citationsarXiv
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation1802.07228v2 · Miles Brundage, Shahar Avin, Jack Clark et al.2018 · 495 citationsarXivon Valency
The Ethics of Advanced AI Assistants2404.16244v2 · Iason Gabriel, Arianna Manzini, Geoff Keeling et al.2024 · 51 citationsarXiv
Holistic Safety and Responsibility Evaluations of Advanced AI Models2404.14068v1 · Laura Weidinger, Joslyn Barnhart, Jenny Brennan et al.2024 · 3 citationsarXiv
Evaluating Frontier Models for Dangerous Capabilities2403.13793v2 · Mary Phuong, Matthew Aitchison, Elliot Catt et al.2024 · 10 citationsarXiv
Beyond Privacy Trade-offs with Structured Transparency2012.08347v2 · Andrew Trask, Emma Bluemke, Teddy Collins et al.2020 · 16 citationsarXiv
Model evaluation for extreme risks2305.15324v2 · Toby Shevlane, Sebastian Farquhar, Ben Garfinkel et al.2023 · 56 citationsarXiv
Randomization Inference beyond the Sharp Null: Bounded Null Hypotheses and Quantiles of Individual Treatment Effects2101.09195v2 · Devin Caughey, Allan Dafoe, Xinran Li et al.2021 · 8 citationsarXiv
Democratising AI: Multiple Meanings, Goals, and Methods2303.12642v3 · Elizabeth Seger, Aviv Ovadya, Ben Garfinkel et al.2023 · 76 citationsarXiv
International Institutions for Advanced AI2307.04699v2 · Lewis Ho, Joslyn Barnhart, Robert Trager et al.2023 · 19 citationsarXiv
Between Progress and Potential Impact of AI: the Neglected Dimensions1806.00610v2 · Fernando Mart'inez-Plumed, Shahar Avin, Miles Brundage et al.2018 · 1 citationarXiv
Forecasting AI Progress: Evidence from a Survey of Machine Learning Researchers2206.04132v1 · Baobao Zhang, Noemi Dreksler, Markus Anderljung et al.2022 · 21 citationsarXiv
Normative Disagreement as a Challenge for Cooperative AI2111.13872v1 · Julian Stastny, Maxime Rich'e, Alexander Lyzhov et al.2021 · 0 citationsarXiv
Cooperative AI: machines must learn to find common ground.33947992 · Dafoe, Allan, Bachrach, Yoram, Hadfield, Gillian et al.2021 · 235 citationsNature. 2021;593(7857):33-36
Engines of Power: Electricity, AI, and General-Purpose Military Transformations2106.04338v1 · Jeffrey Ding, Allan Dafoe2021 · 23 citationsarXiv
The Logic of Strategic Assets: From Oil to Artificial Intelligence2001.03246v2 · Jeffrey Ding, Allan Dafoe2020 · 43 citationsarXiv
Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences2104.07237v2 · Remco Zwetsloot, Baobao Zhang, Noemi Dreksler et al.2021 · 6 citationsarXiv
Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers2105.02117v1 · Baobao Zhang, Markus Anderljung, Lauren Kahn et al.2021 · 84 citationsarXiv
The biosecurity benefits of genetic engineering attribution.33293537 · Lewis, Gregory, Jordan, Jacob L, Relman, David A et al.2021 · 51 citationsNature communications. 2020;11(1):6294
Open Problems in Cooperative AI2012.08630v1 · Allan Dafoe, Edward Hughes, Yoram Bachrach et al.2020 · 15 citationsarXiv
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims2004.07213v2 · Miles Brundage, Shahar Avin, Jasmine Wang et al.2020 · 301 citationsarXiv
A study of the impact of data sharing on article citations using journal policies as a natural experiment.31851689 · Christensen, Garret, Dafoe, Allan, Miguel, Edward et al.2020 · 116 citationsPloS one. 2019;14(12):e0225883
The Windfall Clause: Distributing the Benefits of AI for the Common Good1912.11595v2 · Cullen O'Keefe, Peter Cihon, Ben Garfinkel et al.2019 · 5 citationsarXiv
The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?2001.00463v2 · Toby Shevlane, Allan Dafoe2019 · 8 citationsarXiv
U.S. Public Opinion on the Governance of Artificial Intelligence1912.12835v1 · Baobao Zhang, Allan Dafoe2019 · 90 citationsarXiv
Fitness reduction for uncooperative fig wasps through reduced offspring size: a third component of host sanctions.27859079 · Jandér, K C, Dafoe, A, Herre, E A2018 · 22 citationsEcology. 2016;97(9):2491-2500
When Will AI Exceed Human Performance? Evidence from AI Experts1705.08807v3 · Katja Grace, John Salvatier, Allan Dafoe et al.2017 · 222 citationsarXivon Valency
Beyond the Sharp Null: Randomization Inference, Bounded Null Hypotheses, and Confidence Intervals for Maximum Effects1709.07339v1 · Devin Caughey, Allan Dafoe, Luke Miratrix2017 · 10 citationsarXiv
SCIENTIFIC STANDARDS. Promoting an open research culture.26113702 · Nosek, B A, Alter, G, Banks, G C et al.2015 · 2,856 citationsScience (New York, N.Y.). 2015;348(6242):1422-5
Career total: 435 works. 35 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.