SF
Stanislav Fort
cs.LGstat.MLcs.CVcs.AIcs.CLcs.CRcs.NEastro-ph.IMcs.CYastro-ph.HE
On Valency
published · living versionsW_34tet3c7·v1 · currentpublished
Constitutional AI: Harmlessness from AI Feedback
with Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell +46
1 version
Preprints & journals
36 papers in the corpus · 2015–2026HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models2607.27030v1 · Petr Simecek, Elnaz Babayeva, Jiri Balhar et al.2026 · 0 citationsarXiv
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness2510.06790v3 · Tavish McDonald, Bo Lei, Stanislav Fort et al.2025 · 0 citationsICLR 2026
Solving adversarial examples requires solving exponential misalignment2603.03507v2 · Alessandro Salvatore, Stanislav Fort, Surya Ganguli2026 · 0 citationsarXiv
Representations of Text and Images Align From Layer One2601.08017v1 · Evzen Wybitul, Javier Rando, Florian Tramer et al.2026 · 0 citationsarXiv
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models2502.07753v1 · Stanislav Fort, Jonathan Whitaker2025 · 0 citationsarXiv
A Note on Implementation Errors in Recent Adaptive Attacks Against Multi-Resolution Self-Ensembles2501.14496v1 · Stanislav Fort2025 · 0 citationsarXiv
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness2408.05446v1 · Stanislav Fort, Balaji Lakshminarayanan2024 · 1 citationarXiv
Scaling Laws for Adversarial Attacks on Language Model Activations2312.02780v1 · Stanislav Fort2023 · 0 citationsarXiv
Multi-attacks: Many images $+$ the same adversarial attack $\to$ many target labels2308.03792v1 · Stanislav Fort2023 · 0 citationsarXiv
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
What does a deep neural network confidently perceive? The effective dimension of high certainty class manifolds and their low confidence boundaries2210.05546v1 · Stanislav Fort, Ekin Dogus Cubuk, Surya Ganguli et al.2022 · 0 citationsarXiv
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error2105.13343v2 · Stanislav Fort, Andrew Brock, Razvan Pascanu et al.2021 · 7 citationsarXiv
How many degrees of freedom do we need to train deep networks: a loss landscape perspective2107.05802v2 · Brett W. Larsen, Stanislav Fort, Nic Becker et al.2021 · 2 citationsarXiv
Adversarial vulnerability of powerful near out-of-distribution detection2201.07012v1 · Stanislav Fort2022 · 1 citationarXiv
Training independent subnetworks for robust prediction2010.06610v2 · Marton Havasi, Rodolphe Jenatton, Stanislav Fort et al.2020 · 28 citationsarXiv
Exploring the Limits of Out-of-Distribution Detection2106.03004v3 · Stanislav Fort, Jie Ren, Balaji Lakshminarayanan2021 · 107 citationsarXiv
A Simple Fix to Mahalanobis Distance for Improving Near-OOD Detection2106.09022v1 · Jie Ren, Stanislav Fort, Jeremiah Liu et al.2021 · 70 citationsarXiv
Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes2104.11044v2 · James Lucas, Juhan Bae, Michael R. Zhang et al.2021 · 7 citationsarXiv
Identifying charged particle background events in X-ray imaging detectors with novel machine learning algorithms2012.01463v1 · D. R. Wilkins, S. W. Allen, E. D. Miller et al.2020 · 4 citationsProc. SPIE, 2020, 11444, 308
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel2010.15110v1 · Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul et al.2020 · 42 citationsarXiv
Deep Ensembles: A Loss Landscape Perspective1912.02757v2 · Stanislav Fort, Huiyi Hu, Balaji Lakshminarayanan2019 · 336 citationsarXiv
Stiffness: A New Perspective on Generalization in Neural Networks1901.09491v3 · Stanislav Fort, Pawe\l Krzysztof Nowak, Stanislaw Jastrzebski et al.2019 · 52 citationsarXiv
The Break-Even Point on Optimization Trajectories of Deep Neural Networks2002.09572v1 · Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort et al.2020 · 41 citationsarXiv
Emergent properties of the local geometry of neural loss landscapes1910.05929v1 · Stanislav Fort, Surya Ganguli2019 · 15 citationsarXiv
Large Scale Structure of Neural Network Loss Landscapes1906.04724v1 · Stanislav Fort, Stanislaw Jastrzebski2019 · 46 citationsarXiv
Adaptive Quantum State Tomography with Neural Networks1812.06693v1 · Yihui Quek, Stanislav Fort, Hui Khoon Ng2018 · 95 citationsarXiv
The Goldilocks zone: Towards better understanding of neural network loss landscapes1807.02581v2 · Stanislav Fort, Adam Scherlis2018 · 27 citationsarXiv
The Athena WFI Science Products Module1808.02883v1 · David N. Burrows, Steven Allen, Marshall Bautz et al.2018 · 6 citationsarXiv
Towards understanding feedback from supermassive black holes using convolutional neural networks1712.00523v1 · Stanislav Fort2017 · 1 citationarXiv
Gaussian Prototypical Networks for Few-Shot Learning on Omniglot1708.02735v1 · Stanislav Fort2017 · 60 citationsarXiv
Discovery of Gamma-ray Pulsations from the Transitional Redback PSR J1227-48531502.06862v2 · T. J. Johnson, P. S. Ray, J. Roy et al.2015 · 49 citations2015, ApJ, 806, 91
Career total: 43 works. 36 are in this corpus.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.