AM
Austin Meek
cs.AIcs.LGcs.CLcs.CYbiomarkersdictionary learningDisease Models, AnimalEEGElectroencephalographyepilepsy
On Valency
published · living versionsW_5uruupn8·v1 · currentpublished
Auditing language models for hidden objectives
with Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey +30
1 version
Preprints & journals
12 papers in the corpus · 2023–2026On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness2607.29062v1 · Matthew Nguyen, Kyle Cox, Austin Meek et al.2026 · 0 citationsarXiv
Interpretable EEG biomarkers for neurological disease models in mice using bag-of-waves classifiers.41780177 · Isabel Cano Achuri, Maria, Kay Lara, Montana, Abed Rabbo, Khalil et al.2026 · 0 citationsJournal of neural engineering. 2026;23(3)
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results2606.14516v1 · Jan Batzner, Sree Harsha Nelaturu, Damian Stachura et al.2026 · 0 citationsarXiv
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing2606.00033v1 · Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi et al.2026 · 0 citationsarXiv
Convolutional Monge Mapping between EEG Datasets to Support Independent Component Labeling2509.01721v2 · Austin Meek, Carlos H. Mendoza-Cardenas, Austin J. Brockmeier2025 · 0 citationsarXiv
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity2510.27378v2 · Austin Meek, Eitan Sprejer, Iv'an Arcuschin et al.2025 · 0 citationsarXiv
Interpretable EEG Biomarkers for Neurological Disease Models in Mice Using Bag-of-Waves Classifiers10.1101/2025.08.14.670397v1 · Achuri, M. I. C., Lara, M. K., Rabbo, K. A. et al.2025 · 0 citationsbioRxiv
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders2410.06981v4 · Michael Lan, Philip Torr, Austin Meek et al.2024 · 2 citationsarXiv
Auditing language models for hidden objectives2503.10965v2 · Samuel Marks, Johannes Treutlein, Trenton Bricken et al.2025 · 4 citationsarXivon Valency
Inducing Human-like Biases in Moral Reasoning Language Models2411.15386v1 · Artem Karpov, Seong Hah Cho, Austin Meek et al.2024 · 0 citationsarXiv
Understanding and Controlling a Maze-Solving Policy Network2310.08043v1 · Ulisse Mini, Peli Grietzer, Mrinank Sharma et al.2023 · 0 citationsarXiv
Career total: 18 works. 12 are in this corpus.Showing the 11 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.