SF

Stanislav Fort

cs.LGstat.MLcs.CVcs.AIcs.CLcs.CRcs.NEastro-ph.IMcs.CYastro-ph.HE

On Valency

published · living versions
W_34tet3c7·v1 · currentpublished
Constitutional AI: Harmlessness from AI Feedback
with Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell +46
1 version

Preprints & journals

36 papers in the corpus · 2015–2026
HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models2607.27030v1 · Petr Simecek, Elnaz Babayeva, Jiri Balhar et al.2026 · 0 citationsarXiv
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness2510.06790v3 · Tavish McDonald, Bo Lei, Stanislav Fort et al.2025 · 0 citationsICLR 2026
Solving adversarial examples requires solving exponential misalignment2603.03507v2 · Alessandro Salvatore, Stanislav Fort, Surya Ganguli2026 · 0 citationsarXiv
Representations of Text and Images Align From Layer One2601.08017v1 · Evzen Wybitul, Javier Rando, Florian Tramer et al.2026 · 0 citationsarXiv
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness2408.05446v1 · Stanislav Fort, Balaji Lakshminarayanan2024 · 1 citationarXiv
Constitutional AI: Harmlessness from AI Feedback2212.08073v1 · Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.2022 · 322 citationsarXivon Valency
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2209.07858v2 · Deep Ganguli, Liane Lovitt, Jackson Kernion et al.2022 · 119 citationsarXivon Valency
Language Models (Mostly) Know What They Know2207.05221v4 · Saurav Kadavath, Tom Conerly, Amanda Askell et al.2022 · 169 citationsarXivon Valency
Measuring Progress on Scalable Oversight for Large Language Models2211.03540v2 · Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez et al.2022 · 34 citationsarXiv
Predictability and Surprise in Large Generative Models2202.07785v2 · Deep Ganguli, Danny Hernandez, Liane Lovitt et al.2022 · 204 citationsarXiv
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback2204.05862v1 · Yuntao Bai, Andy Jones, Kamal Ndousse et al.2022 · 389 citationsarXiv
Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error2105.13343v2 · Stanislav Fort, Andrew Brock, Razvan Pascanu et al.2021 · 7 citationsarXiv
How many degrees of freedom do we need to train deep networks: a loss landscape perspective2107.05802v2 · Brett W. Larsen, Stanislav Fort, Nic Becker et al.2021 · 2 citationsarXiv
Training independent subnetworks for robust prediction2010.06610v2 · Marton Havasi, Rodolphe Jenatton, Stanislav Fort et al.2020 · 28 citationsarXiv
Exploring the Limits of Out-of-Distribution Detection2106.03004v3 · Stanislav Fort, Jie Ren, Balaji Lakshminarayanan2021 · 107 citationsarXiv
A Simple Fix to Mahalanobis Distance for Improving Near-OOD Detection2106.09022v1 · Jie Ren, Stanislav Fort, Jeremiah Liu et al.2021 · 70 citationsarXiv
Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes2104.11044v2 · James Lucas, Juhan Bae, Michael R. Zhang et al.2021 · 7 citationsarXiv
Identifying charged particle background events in X-ray imaging detectors with novel machine learning algorithms2012.01463v1 · D. R. Wilkins, S. W. Allen, E. D. Miller et al.2020 · 4 citationsProc. SPIE, 2020, 11444, 308
Deep Ensembles: A Loss Landscape Perspective1912.02757v2 · Stanislav Fort, Huiyi Hu, Balaji Lakshminarayanan2019 · 336 citationsarXiv
Stiffness: A New Perspective on Generalization in Neural Networks1901.09491v3 · Stanislav Fort, Pawe\l Krzysztof Nowak, Stanislaw Jastrzebski et al.2019 · 52 citationsarXiv
The Break-Even Point on Optimization Trajectories of Deep Neural Networks2002.09572v1 · Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort et al.2020 · 41 citationsarXiv
Emergent properties of the local geometry of neural loss landscapes1910.05929v1 · Stanislav Fort, Surya Ganguli2019 · 15 citationsarXiv
Large Scale Structure of Neural Network Loss Landscapes1906.04724v1 · Stanislav Fort, Stanislaw Jastrzebski2019 · 46 citationsarXiv
Adaptive Quantum State Tomography with Neural Networks1812.06693v1 · Yihui Quek, Stanislav Fort, Hui Khoon Ng2018 · 95 citationsarXiv
The Athena WFI Science Products Module1808.02883v1 · David N. Burrows, Steven Allen, Marshall Bautz et al.2018 · 6 citationsarXiv
Discovery of Gamma-ray Pulsations from the Transitional Redback PSR J1227-48531502.06862v2 · T. J. Johnson, P. S. Ray, J. Roy et al.2015 · 49 citations2015, ApJ, 806, 91
Career total: 43 works. 36 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.