LM

Leon Maksin

cs.AIcs.CLcs.LG

On Valency

published · living versions
W_6cysh3yv·v1 · currentpublished
PaperBench: Evaluating AI's Ability to Replicate AI Research
with Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung +8
1 version

Preprints & journals

4 papers in the corpus · 2024–2026
Predicting LLM Safety Before Release by Simulating Deployment2607.07184v1 · Marcus Williams, Hannah Sheahan, Cameron Raymond et al.2026 · 0 citationsarXiv
OpenAI GPT-5 System Card2601.03267v2 · Aaditya Singh, Adam Fry, Adam Perelman et al.2025 · 18 citationsarXiv
OpenAI o1 System Card2412.16720v2 · OpenAI: Aaron Jaech, Adam Kalai, Adam Lerer et al.2024 · 44 citationsarXiv
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering2410.07095v6 · Jun Shern Chan, Neil Chowdhury, Oliver Jaffe et al.2024 · 9 citationsarXiv
Career total: 4 works. 4 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.