SF

Sebastian Farquhar

cs.LGstat.MLcs.AIcs.CLcs.CRcs.CVcs.CYcs.ITcs.MAeess.IV

On Valency

published · living versions
W_y3274kcu·v1 · currentpublished
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
with Miles Brundage, Shahar Avin, Jack Clark, Helen Toner +21
1 version

Preprints & journals

36 papers in the corpus · 2018–2026
GDM AI Control Roadmap2607.13087v1 · Mary Phuong, Erik Jenner, Laurent Simon et al.2026 · 0 citationsarXiv
Realistic honeypot evaluations for scheming propensity2605.29729v2 · Victoria Krakovna, David Lindner, Lewis Ho et al.2026 · 0 citationsarXiv
Gram: Assessing sabotage propensities via automated alignment auditing2605.30322v1 · David Lindner, Victoria Krakovna, Sebastian Farquhar2026 · 0 citationsarXiv
Practical challenges of control monitoring in frontier AI deployments2512.22154v1 · David Lindner, Charlie Griffin, Tomek Korbak et al.2025 · 0 citationsarXiv
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking2501.13011v2 · Sebastian Farquhar, Vikrant Varma, David Lindner et al.2025 · 0 citationsarXiv
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation1802.07228v2 · Miles Brundage, Shahar Avin, Jack Clark et al.2018 · 495 citationsarXivon Valency
Detecting hallucinations in large language models using semantic entropy.38898292 · Farquhar, Sebastian, Kossen, Jannik, Kuhn, Lorenz et al.2024 · 852 citationsNature. 2024;630(8017):625-630
Holistic Safety and Responsibility Evaluations of Advanced AI Models2404.14068v1 · Laura Weidinger, Joslyn Barnhart, Jenny Brennan et al.2024 · 3 citationsarXiv
Evaluating Frontier Models for Dangerous Capabilities2403.13793v2 · Mary Phuong, Matthew Aitchison, Elliot Catt et al.2024 · 10 citationsarXiv
Challenges with unsupervised LLM knowledge discovery2312.10029v2 · Sebastian Farquhar, Vikrant Varma, Zachary Kenton et al.2023 · 3 citationsarXiv
Tracr: Compiled Transformers as a Laboratory for Interpretability2301.05062v5 · David Lindner, J'anos Kram'ar, Sebastian Farquhar et al.2023 · 7 citationsarXiv
Prioritized training on points that are learnable, worth learning, and not yet learned (workshop version)2107.02565v4 · Soren Mindermann, Muhammed Razzak, Winnie Xu et al.2021 · 0 citationsICML 2021 Workshop on Subset Selection in Machine Learning
Model evaluation for extreme risks2305.15324v2 · Toby Shevlane, Sebastian Farquhar, Ben Garfinkel et al.2023 · 56 citationsarXiv
Stochastic Batch Acquisition: A Simple Baseline for Deep Active Learning2106.12059v3 · Andreas Kirsch, Sebastian Farquhar, Parmida Atighehchian et al.2021 · 3 citationsarXiv
Prediction-Oriented Bayesian Active Learning2304.08151v1 · Freddie Bickford Smith, Andreas Kirsch, Sebastian Farquhar et al.2023 · 6 citationsarXiv
Do Bayesian Neural Networks Need To Be Fully Stochastic?2211.06291v2 · Mrinank Sharma, Sebastian Farquhar, Eric Nalisnick et al.2022 · 5 citationsarXiv
CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models2212.07769v2 · Lorenz Kuhn, Yarin Gal, Sebastian Farquhar2022 · 7 citationsarXiv
Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation2202.06881v2 · Jannik Kossen, Sebastian Farquhar, Yarin Gal et al.2022 · 4 citationsarXiv
Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt2206.07137v3 · Soren Mindermann, Jan Brauner, Muhammed Razzak et al.2022 · 19 citationsarXiv
Discovering Agents2208.08345v2 · Zachary Kenton, Ramana Kumar, Sebastian Farquhar et al.2022 · 7 citationsarXiv
Path-Specific Objectives for Safer Agent Incentives2204.10018v1 · Sebastian Farquhar, Ryan Carey, Tom Everitt2022 · 12 citationsarXiv
Prospect Pruning: Finding Trainable Weights at Initialization using Meta-Gradients2202.08132v2 · Milad Alizadeh, Shyam A. Tailor, Luisa M Zintgraf et al.2022 · 11 citationsarXiv
Uncertainty Baselines: Benchmarks for Uncertainty & Robustness in Deep Learning2106.04015v3 · Zachary Nado, Neil Band, Mark Collier et al.2021 · 36 citationsarXiv
Active Testing: Sample-Efficient Model Evaluation2103.05331v2 · Jannik Kossen, Sebastian Farquhar, Yarin Gal et al.2021 · 16 citationsarXiv
Radial Bayesian Neural Networks: Beyond Discrete Support In Large-Scale Bayesian Deep Learning1907.00865v4 · Sebastian Farquhar, Michael Osborne, Yarin Gal2019 · 43 citationsAI Stats, PMLR 108:1352-1362, 2020
On Statistical Bias In Active Learning: How and When To Fix It2101.11665v2 · Sebastian Farquhar, Yarin Gal, Tom Rainforth2021 · 33 citationsarXiv
Single Shot Structured Pruning Before Training2007.00389v1 · Joost van Amersfoort, Milad Alizadeh, Sebastian Farquhar et al.2020 · 14 citationsarXiv
A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks1912.10481v1 · Angelos Filos, Sebastian Farquhar, Aidan N. Gomez et al.2019 · 71 citationsarXiv
Towards Robust Evaluations of Continual Learning1805.09733v3 · Sebastian Farquhar, Yarin Gal2018 · 177 citationsarXiv
A Unifying Bayesian View of Continual Learning1902.06494v1 · Sebastian Farquhar, Yarin Gal2019 · 46 citationsarXiv
Differentially Private Continual Learning1902.06497v1 · Sebastian Farquhar, Yarin Gal2019 · 6 citationsarXiv
Pricing Externalities to Balance Public Risks and Benefits of Research.28767274 · Farquhar, Sebastian, Cotton-Barratt, Owen, Snyder-Beattie, Andrew2018 · 9 citationsHealth security. 2017;15(4):401-408
Career total: 46 works. 36 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.